Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “tensor learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

torch-einshard v1.0

torch-einshard is a Python library for describing local and distributed PyTorch tensor computations with compact, einsum-like notation. Its expressions name logical axes, specify how they are sharded across a PyTorch DeviceMesh, and represent partial reductions. The library automatically performs contractions, permutations, reshaping, splitting, gathering, reduction, reduce-scatter, and repartitioning while preserving autograd. Additional features include sharding-aware FFTs, tensor rolls, halo exchange, sliding windows, 1D–3D convolutions, uneven-shard handling, parameter initialization and gradient management, and cost-based execution planning. It is designed for scientific machine learning and large-model workloads, including tensor-, sequence-, and spatial-parallel MLPs, attention, convolutions, and spectral operations. Compared with manually combining torch.einsum and distributed collectives, torch-einshard expresses both the mathematical operation and data placement in one readable formula. This reduces boilerplate and synchronization errors, keeps forward and backward communication consistent, and allows the library to select optimized collective strategies without changing model code.

Morozov, Dmitriy [Lawrence Berkeley National Labor↗

Estimating Higher-Order Moments Using Symmetric Tensor Decomposition

In this paper, we consider the problem of decomposing higher-order moment tensors, i.e., the sum of symmetric outer products of data vectors. Such a decomposition can be used to estimate the means in a Gaussian mixture model and for other applications in machine learning. The dth-order empirical moment tensor of a set of p observations of n variables is a symmetric d-way tensor. Our goal is to nd a low-rank tensor approximation comprising r $\ll$ p symmetric outer products. The challenge is that forming the empirical moment tensor costs O(pn d ) operations and O(n d ) storage, which may be prohibitively expensive; additionally, the algorithm to compute the low-rank approximation costs O(n d ) per iteration. Our contribution is avoiding formation of the moment tensor, computing the low-rank tensor approximation of the moment tensor implicitly using O(pnr) operations per iteration and no extra memory. This advance opens the door to more applications of higher-order moments since they can now be efficiently computed. We present numerical evidence of the computational savings and show an example of estimating the means for higher-order moments.

97 MATHEMATICS AND COMPUTING↗

Nonnegative canonical tensor decomposition with linear constraints: nnCANDELINC

Abstract There is an emerging interest for tensor factorization applications in big‐data analytics and machine learning. To speed up the factorization of extra‐large datasets, organized in multidimensional arrays (also known as tensors), easy to compute compression‐based tensor representations, such as, Tucker and tensor train formats, are used to approximate the initial large‐tensor. Further, tensor factorization is used to extract latent features that can facilitate discoveries of new mechanisms and signatures hidden in the data, where the explainability of the latent features is of principal importance. Nonnegative tensor factorization extracts latent features that are naturally sparse and parts of the data, which makes them easily interpretable. However, to take into account available domain knowledge and subject matter expertise, often additional constraints need to be imposed, which lead us to canonical decomposition with linear constraints (CANDELINC), a canonical polyadic decomposition with rank deficient factors. In CANDELINC, Tucker compression is used as a preprocessing step, which lead to a larger residual error but to more explainable latent features. Here, we propose a nonnegative CANDELINC (nnCANDELINC) accomplished via a specific nonnegative Tucker decomposition; we refer to as minimal or canonical nonnegative Tucker. We derive several results required to understand the specificity of nnCANDELINC, focusing on the difficulties of preserving the nonnegative rank of a tensor to its Tucker core and comparing the real valued to nonnegative case. Finally, we demonstrate nnCANDELINC performance on synthetic and real‐world examples.

97 MATHEMATICS AND COMPUTING↗

tesuract v.1.0

SAND2021-14370 O Tesuract (tensor surrogate approximation and computation) provides Python tools to build machine learning regressors that include polynomial chaos expansions. The software also offers methods for creation regressors for tensor outputs such as spatially varying fields and using a reduced order modeling methodology. The library is built on and compatible with the scikit-learn API, which allows integration with one of the most ubiquitous machine learning Python libraries out there. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Chowdhary, Kamaljit↗

On the Feasibility of Using Reduced-Precision Tensor Core Operations for Graph Analytics

Today’s data-driven analytics and machine learning workload have been largely driven by the General-PurposeGraphics Processing Units (GPGPUs). To accelerate dense matrix multiplications on the GPUs, Tensor Core Units (TCUs) have been introduced in recent years. In this paper, we study linear-algebra-based and vertex-centric algorithms for various graph kernels on the GPUs with an objective of applying this new hardware feature to graph applications. We identify the potential stages in these graph kernels that can be executed on the Tensor Core Units. In particular, we leverage the reformulation of the reduction and scan operations in terms of matrix multiplication [1]on the TCUs. We demonstrate that executing these operations on the TCUs, available inside different graph kernels, can assist in establishing an end-to-end pipeline on the GPGPUs without depending on hand-tuned external libraries and still can deliver comparable performance for various graph analytics.

Graph algorithms, GPU computing↗

Machine learning interatomic potential for predicting the thermal properties of uranium nitride

We present a combined computational and experimental investigation of the thermal properties of uranium nitride (UN), focusing on the development of a machine learning interatomic potential (MLIP) using the moment tensor potential framework. The MLIP was trained on density functional theory (DFT) data and validated against various quantities including energies, forces, elastic constants, phonon dispersion, and defect formation energies, achieving excellent agreement with DFT calculations, prior experimental results, and our thermal conductivity measurement. The potential was then employed in molecular dynamics simulations to predict key thermal properties such as melting point, thermal expansion, specific heat, and lattice thermal conductivity. To further assess model accuracy, we fabricated a UN sample and performed new thermal conductivity measurements representative of single-crystal properties, which showed strong agreement with the MLIP predictions. This work confirms the reliability and predictive capability of the developed potential for determining the thermal properties of UN.

36 - MATERIALS SCIENCE↗

Discrete generative diffusion models without stochastic differential equations: A tensor network approach

Diffusion models (DMs) are a class of generative machine learning methods that sample a target distribution by transforming samples of a trivial (often Gaussian) distribution using a learned stochastic differential equation. In standard DMs, this is done by learning a “score function” that reverses the effect of adding diffusive noise to the distribution of interest. Here we consider the generalisation of DMs to lattice systems with discrete degrees of freedom, and where noise is added via Markov chain jump dynamics. We show how to use tensor networks (TNs) to efficiently define and sample such “discrete diffusion models” (DDMs) without explicitly having to solve a stochastic differential equation. We show the following: (i) by parametrising the data and evolution operators as TNs, the denoising dynamics can be represented exactly; (ii) the auto-regressive nature of TNs allows to generate samples efficiently and without bias; (iii) for sampling Boltzmann-like distributions, TNs allow to construct an efficient learning scheme that integrates well with Monte Carlo. We illustrate this approach to study the equilibrium of two models with non-trivial thermodynamics, the d = 1 constrained Fredkin chain and the d = 2 Ising model. Published by the American Physical Society 2025

Causer, Luke (ORCID:0000000194243473)↗

HST Inertia Tensor Optimization for Two-Gyro Science

The power point presentation includes an overview of the Hubble Space Telescope, pointing control hardware peculiarities, the two-gyro science (TGS) control system, TGS modifications to preserved hardware lifetime, chasing-down disturbance torques less than or equal to 0.002 Nm, inertia tensor optimization, and a summary with lessons learned.

Clapp, Brian↗

Static versioning in the polyhedral model

An approach is presented to enhancing the optimization process in a polyhedral compiler by introducing compile-time versioning, i.e., the production of several versions of optimized code under varying assumptions on its run-time parameters. We illustrate this process by enabling versioning in the polyhedral processor placement pass. We propose an efficient code generation method and validate that versioning can be useful in a polyhedral compiler by performing benchmarking on a small set of deep learning layers defined for dynamically-sized tensors.

Meister, Benoit J.↗

Dimension Reduction and Redundancy Removal through Successive Schmidt Decompositions

Quantum computers are believed to have the ability to process huge data sizes, which can be seen in machine learning applications. In these applications, the data, in general, are classical. Therefore, to process them on a quantum computer, there is a need for efficient methods that can be used to map classical data on quantum states in a concise manner. On the other hand, to verify the results of quantum computers and study quantum algorithms, we need to be able to approximate quantum operations into forms that are easier to simulate on classical computers with some errors. Motivated by these needs, in this paper, we study the approximation of matrices and vectors by using their tensor products obtained through successive Schmidt decompositions. We show that data with distributions such as uniform, Poisson, exponential, or similar to these distributions can be approximated by using only a few terms, which can be easily mapped onto quantum circuits. The examples include random data with different distributions, the Gram matrices of iris flower, handwritten digits, 20newsgroup, and labeled faces in the wild. Similarly, some quantum operations, such as quantum Fourier transform and variational quantum circuits with a small depth, may also be approximated with a few terms that are easier to simulate on classical computers. Furthermore, we show how the method can be used to simplify quantum Hamiltonians: In particular, we show the application to randomly generated transverse field Ising model Hamiltonians. The reduced Hamiltonians can be mapped into quantum circuits easily and, therefore, can be simulated more efficiently.

97 MATHEMATICS AND COMPUTING↗

ExaTN: Scalable GPU-Accelerated High-Performance Processing of General Tensor Networks at Exascale

We present ExaTN (Exascale Tensor Networks), a scalable GPU-accelerated C++ library which can express and process tensor networks on shared- as well as distributed-memory high-performance computing platforms, including those equipped with GPU accelerators. Specifically, ExaTN provides the ability to build, transform, and numerically evaluate tensor networks with arbitrary graph structures and complexity. It also provides algorithmic primitives for the optimization of tensor factors inside a given tensor network in order to find an extremum of a chosen tensor network functional, which is one of the key numerical procedures in quantum many-body theory and quantum-inspired machine learning. Numerical primitives exposed by ExaTN provide the foundation for composing rather complex tensor network algorithms. We enumerate multiple application domains which can benefit from the capabilities of our library, including condensed matter physics, quantum chemistry, quantum circuit simulations, as well as quantum and classical machine learning, for some of which we provide preliminary demonstrations and performance benchmarks just to emphasize a broad utility of our library.

97 MATHEMATICS AND COMPUTING↗

CP decomposition for tensors via alternating least squares with QR decomposition

The CP tensor decomposition is used in applications such as machine learning and signal processing to discover latent low-rank structure in multidimensional data. Computing a CP decomposition via an alternating least squares (ALS) method reduces the problem to several linear least squares problems. The standard way to solve these linear least squares subproblems is to use the normal equations, which inherit special tensor structure that can be exploited for computational efficiency. However, the normal equations are sensitive to numerical ill-conditioning, which can compromise the results of the decomposition. In this paper, we develop versions of the CP-ALS algorithm using the QR decomposition and the singular value decomposition, which are more numerically stable than the normal equations, to solve the linear least squares problems. Our algorithms utilize the tensor structure of the CP-ALS subproblems efficiently, have the same complexity as the standard CP-ALS algorithm when the input is dense and the rank is small, and are shown via examples to produce more stable results when ill-conditioning is present. Our MATLAB implementation achieves the same running time as the standard algorithm for small ranks, and we show that the new methods can obtain lower approximation error.

97 MATHEMATICS AND COMPUTING↗

A Machine Learning Model for Solar Sail Shape Reconstruction Using Flight Data

Solar sail deformation leads to disturbance torques from solar radiation pressure, driving performance requirements for momentum management systems. For the Solar Cruiser technology demonstrator mission, we have developed a model leveraging neural network-based machine learning to derive sail shape characteristics. The model uses torque and attitude telemetry simulated from a reduced-order tensor model of the deformed sail mesh over a characterization sequence. The machine learning model predicts sail boom deflection with comparable accuracy to that of an onboard context camera. This model can discover sail shape with no additional mass or data downlink requirements, allowing for validation of sail force modeling assumptions using in flight data. The results from the project hold promise for the further implementation of machine learning techniques in solar sail telemetry analysis and control.

solar sail↗

A Machine Learning Model for Solar Sail Shape Reconstruction Using Flight Data

Solar sail deformation leads to disturbance torques from solar radiation pressure, driving performance requirements for momentum management systems. For the Solar Cruiser technology demonstrator mission, we have developed a model leveraging neural network-based machine learning to derive sail shape characteristics. The model uses torque and attitude telemetry simulated from a reduced-order tensor model of the deformed sail mesh over a characterization sequence. The machine learning model predicts sail boom deflection with comparable accuracy to that of an onboard context camera. This model can discover sail shape with no additional mass or data downlink requirements, allowing for validation of sail force modeling assumptions using in flight data. The results from the project hold promise for the further implementation of machine learning techniques in solar sail telemetry analysis and control.

solar sail↗

AXI4MLIR: User-Driven Automatic Host Code Generation for Custom AXI-Based Accelerators

Tensor algebra operations represent an important class of algorithms used across many applications, including machine learning, scientific computing, and data analytics. As a result, the efficient generation of custom accelerators for tensor operations has received increased attention. Previous efforts have produced automated tools enabling users to prototype and explore optimized accelerators. However, little effort has been focused on the host-accelerator interaction in these tools. Efficient use of hardware accelerators requires knowledge about the accelerator's capabilities (operations, data formats, and opcode support), the host CPU microarchitecture (e.g., memory hierarchy), the host-accelerator interface, and the application's features (which code regions should be mapped onto an accelerator). Manually rewriting the original applications to facilitate improved custom accelerator mapping is an error-prone and time-consuming endeavor. To cope with this, we propose AXI4MLIR, a new framework to automatically generate and optimize the communication between the host CPU and arbitrary accelerators that implement linear algebra algorithms. AXI4MLIR extends the MLIR compiler framework to automatically generate efficient host-accelerator driver code for accelerators with AXI-based interfaces. Our compiler extensions enable automatic driver code generation while carefully considering the host's memory hierarchy and target accelerator features. To demonstrate the flexibility and utility of AXI4MLIR, we test it with diverse use cases that include different types of accelerators, tiling scenarios, and dataflow schemes. We compare our experimental results to manual implementations of host-accelerator driver code and find that our approach can reduce CPU cache references by 56% and deliver up to a 1.65x speedup.

Bohm Agostini, Nicolas↗

An end-to-end trainable hybrid classical-quantum classifier

Abstract We introduce a hybrid model combining a quantum-inspired tensor network and a variational quantum circuit to perform supervised learning tasks. This architecture allows for the classical and quantum parts of the model to be trained simultaneously, providing an end-to-end training framework. We show that compared to the principal component analysis, a tensor network based on the matrix product state with low bond dimensions performs better as a feature extractor for the input data of the variational quantum circuit in the binary and ternary classification of MNIST and Fashion-MNIST datasets. The architecture is highly adaptable and the classical-quantum boundary can be adjusted according to the availability of the quantum resource by exploiting the correspondence between tensor networks and quantum circuits.

97 MATHEMATICS AND COMPUTING↗

Union: A Unified HW-SW Co-Design Ecosystem in MLIR for Evaluating Tensor Operationson Spatial Accelerators

To meet the extreme compute demands for deep learning across commercial and scientific applications, dataflow accelerators are becoming increasingly popular. While these“domain-specific” accelerators are not fully programmable like CPUs and GPUs, they retain varying levels of flexibility with respect to data orchestration, i.e., dataflow and tiling optimizations to enhance efficiency. There are several challenges when designing new algorithms and mapping approaches to execute the algorithms for a target problem on new hardware. Previous works have addressed these challenges individually. To address this challenge as a whole, in this work, we present an HW-SW co-design ecosystem for spatial accelerators called Union within the popular MLIR compiler infrastructure. Our framework allows exploring different algorithms and their mappings on several accelerator cost models. Union also includes a plug-and-play library of accelerator cost models and mappers which can easily be extended. The algorithms and accelerator cost models are connected via a novel mapping abstraction that captures the map space of spatial accelerators which can be systematically pruned based on constraints from the hardware, workload, and mapper. We demonstrate the value of Union for the community with several case studies which examine offloading different tensor operations (CONV/GEMM/Tensor Contraction) on diverse accelerator architectures using different mapping schemes.

Jeong, Geonhwa↗