Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Matrix”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Communication Lower Bounds and Optimal Algorithms for Multiple Tensor-Times-Matrix Computation

Multiple tensor-times-matrix (Multi-TTM) is a key computation in algorithms for computing and operating with the Tucker tensor decomposition, which is frequently used in multidimensional data analysis. Here, we establish communication lower bounds that determine how much data movement is required (under mild conditions) to perform the Multi-TTM computation in parallel. The crux of the proof relies on analytically solving a constrained, nonlinear optimization problem. We also present a parallel algorithm to perform this computation that organizes the processors into a logical grid with twice as many modes as the input tensor. We show that, with correct choices of grid dimensions, the communication cost of the algorithm attains the lower bounds and is therefore communication optimal. Finally, we show that our algorithm can significantly reduce communication compared to the straightforward approach of expressing the computation as a sequence of tensor-times-matrix operations when the input and output tensors vary greatly in size.

HBL-inequalities↗

IEEE 123-Bus System A Matrix

System component matrix for IEEE 123 bus distribution system for dynamic simulations.

Sahu, Vibhuti [Oak Ridge National Laboratory] (ORC↗

Evaluating Spatial Accelerator Architectures with Tiled Matrix-Matrix Multiplication

There is a growing interest in custom spatial accelerators for machine learning applications. These accelerators employ a spatial array of processing elements (PEs) interacting via custom buffer hierarchies and networks-on-chip. The efficiency of these accelerators comes from employing optimized dataflow (i.e., spatial/temporal partitioning of data across the PEs and fine-grained scheduling) strategies to optimize data reuse. The focus of this work is to evaluate these accelerator architectures using a tiled general matrix-matrix multiplication (GEMM) kernel. To do so, we develop a framework that finds optimized mappings (dataflow and tile sizes) for a tiled GEMM for a given spatial accelerator and workload combination, leveraging an analytical cost model for runtime and energy. Our evaluations over five spatial accelerators demonstrate that the tiled GEMM mappings systematically generated by our framework achieve high performance on various GEMM workloads and accelerators.

43 PARTICLE ACCELERATORS↗

R-matrix School 2025: Introduction to R-matrix Theory

This technical memo serves as a lecture material for the R-matrix school 2025 and distributed among the participants. The manuscript discusses in details the introduction to the R-matrix theory including its algorithm to calculate reaction cross sections and derivation.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Architecture-Aware Models of AI Engines for High-Performance Matrix Matrix Multiplication

The AI Engine (AIE) architecture, available in systems from mobile SoCs to server-class FPGAs, aims to efficiently execute AI/ML tasks through a two-dimensional array of compute tiles. Previous work on AIEs has explored different approaches to mapping computation across spatial arrays, but the compute kernel running on each tile has not been the focus. Additionally, the AIE-ML architecture introduces memory tiles and omits programmable logic, requiring new approaches to staging and moving data throughout the array. In this work we update analytical models developed for CPUs to produce the design of high performance kernels while introducing new model considerations such as memory structure, throughput, and latency as required by the AIE hardware. We evaluate our models by developing AIE-ML kernels for matrix multiplication in low-precision data types showing performance up to 95% of compute peak for the kernel when data resides in local memory and above 90% of compute peak when data resides in main memory.

Binder, Elliott D. [Carnegie Mellon University, Pi↗

Numerical Simulation of Vibrational Sum Frequency Generation Intensity for Non-Centrosymmetric Domains Interspersed in an Amorphous Matrix: A Case Study for Cellulose in Plant Cell Wall

Vibrational sum frequency generation (SFG) spectroscopy can specifically probe molecular species non-centrosymmetrically arranged in a centrosymmetric or isotropic medium. This capability has been extensively utilized to detect and study molecular species present at the two-dimensional (2D) interface at which the centrosymmetry or isotropy of bulk phases is naturally broken. The same principle has been demonstrated to be very effective for the selective detection of non-centrosymmetric crystalline nanodomains interspersed in three-dimensional (3D) amorphous phases. However, the full spectral interpretation of SFG features has been difficult due to the complexity associated with the theoretical calculation of SFG responses of such 3D systems. This paper describes a numerical method to predict the relative SFG intensities of non-centrosymmetric nanodomains in 3D systems as functions of their size and concentration as well as their assembly patterns, i.e., the distributions of tilt, azimuth, and rotation angles with respect to the lab coordinate. We applied the developed method to predict changes in the CH and OH stretch modes characteristic to crystalline cellulose microfibrils distributed with various orders, which are relevant to plant cell wall structures. As a result, the same algorithm can also be applied to any SFG-active nanodomains interspersed in 3D amorphous matrices.

36 MATERIALS SCIENCE↗

A thermodynamics-based damage model for the non-linear mechanical behavior of SiC/SiC ceramic matrix composites in irradiation and thermal environments

A damage model is developed and validated with experimental data for the non-linear mechanical behavior of SiC/SiC composite materials in nuclear applications. Cyclic thermal and mechanical loading associated with neutron irradiation effects of these composites leads to wide-spread and progressive micro-cracking that leads to loss of thermal conductivity and further enhancement of thermo-mechanical damage. A physics-based model of wide-spread micro-cracking is developed within the thermodynamic framework of continuum damage mechanics. Evolution equations for damage parameters that describe the growth of continuum damage are developed, where the material variables are obtained from experiments. The model novelty is in coupling mechanical, thermal, and irradiation damage through a consistent thermodynamic framework, including loss of thermal conductivity due to the evolution of mechanically induced micro-cracks. A number of thermo-mechanical experiments were conducted to confirm model assumptions. The model is shown to be validated with out-of-pile experiments, and then implemented using commercial finite element code COMSOL to the fuel cladding problem with normal and high radiation dose cases.

Materials Science↗