Engineering PapersSearch

SEARCH · Engineering Papers

Results for “tensor times same vector”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Scalable Tensor Methods for Nonuniform Hypergraphs

While multilinear algebra appears natural for studying the multiway interactions modeled by hypergraphs, tensor methods for general hypergraphs have been stymied by theoretical and practical barriers. A recently proposed adjacency tensor is applicable to nonuniform hypergraphs, but is prohibitively costly to form and analyze in practice. We develop tensor times same vector (TTSV) algorithms for this tensor which improve complexity from $O(n^r)$ to a low-degree polynomial in $r$, where $n$ is the number of vertices and $r$ is the maximum hyperedge size. Our algorithms are implicit, avoiding formation of the order $r$ adjacency tensor. Here, we demonstrate the flexibility and utility of our approach in practice by developing tensor-based hypergraph centrality and clustering algorithms. We also show these tensor measures offer complementary information to analogous graph-reduction approaches on data, and are also able to detect higher-order structure that many existing matrix-based approaches provably cannot.

97 MATHEMATICS AND COMPUTING

Search for charged-lepton flavour violation in top quark interactions with an up-type quark, a muon, and a $τ$ lepton in proton-proton collisions at $\sqrt{s}$ = 13 TeV

A search for charged-lepton flavour violation (CLFV) in top quark (t) production and decay is presented. The search uses proton-proton collision data corresponding to 138 fb$^{-1}$ collected with the CMS experiment at $\sqrt{s}$ = 13 TeV. The signal consists of the production of a single top quark via a CLFV interaction or top quark pair production followed by a CLFV decay. The analysis selects events containing a hadronically decaying $τ$ lepton and a muon of opposite electric charge, as well as at least three jets, one of which is identified as originating from the fragmentation of a bottom quark. Machine learning classification techniques are used to distinguish signal from standard model background events. The results of this search are consistent with the standard model expectations. The upper limits at 95% confidence level on the branching fraction $\mathcal{B}$ for CLFV top quark decays to a muon, a $τ$ lepton, and an up or a charm quark are set at $\mathcal{B}$(t $\to$ $μτ$u) $\lt$ (0.04, 0.08, and 0.12) $\times$ 10$^{-6}$, and $\mathcal{B}$(t $\to$ $μτ$c) $\lt$ (0.81, 1.71, and 2.05) $\times$ 10$^{-6}$ for scalar, vector, and tensor-like operators, respectively.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Decoding the B → K ν ν excess at Belle II: Kinematics, operators, and masses

An excess in the branching fraction for B + → K + ν ν recently measured at Belle II may be a hint of new physics. We perform thorough likelihood analyses for different new physics scenarios such as B → K X with a new invisible particle X , or B → K χ χ through a scalar, vector, or tensor current with χ being a new invisible particle or a neutrino. We find that vector-current three-body decay with m X ≃ 0.6 GeV—which may be dark matter—is most favored, while two-body decay with m X ≃ 2 GeV is also competitive. The best-fit branching fractions for the scalar and tensor cases are a few times larger than for the two-body and vector cases. Past measurements provide further discrimination, although the best-fit parameters stay similar. Published by the American Physical Society 2024

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

NEML2: A High Performance Library for Constitutive Modeling

NEML2, the New Engineering Material model Library, version 2, is an offshoot of NEML, an earlier material modeling code developed at Argonne National Laboratory. NEML2 extends the key philosophy of its predecessor, i.e., material models are flexible, modular, and can be built from smaller blocks. It also provides modern features that do not exist in the framework of its predecessor such as material model vectorization, automatic differentiation, device-portable just-in-time compilation, operator fusion, lazy tensor evaluation, etc. Moreover, NEML2 can seamlessly integrate with the popular machine learning package PyTorch to take advantage of modern and fast-growing machine learning techniques. In this fiscal year, the development of core library features and capabilities are complete. The purpose of this report is not to serve as a verbatim copy of the software API reference (which is available online at https://reverendbedford.github.io/neml2/). Instead, this report documents the motivation, implementation, design choices, and usage of each core capability as well as their applications in solving practical engineering problems. This report is compiled based on the NEML2 major release 2.0.0.

36 MATERIALS SCIENCE

Far-from-equilibrium slow modes and momentum anisotropy in an expanding plasma

The momentum distribution of particle production in heavy-ion collisions encodes information about thermalization processes in the early-stage quark-gluon plasma. We use kinetic theory to study the far-from-equilibrium evolution of an expanding plasma with an anisotropic momentum-space distribution. We identify slow and fast degrees of freedom in the far-from-equilibrium plasma from the evolution of moments of this distribution. At late times, the slow modes correspond to hydrodynamic degrees of freedom and are naturally gapped from the fast modes by the inverse of the relaxation time, τ R − 1 . At early times, however, there are an infinite number of slow modes with a gap inversely proportional to time, τ − 1 . From the evolution of the slow modes we generalize the paradigm of the far-from-equilibrium attractor to vector and tensor components of the energy-momentum tensor, and even to higher moments of the distribution function that are not part of the hydrodynamic evolution. We predict that initial-state momentum anisotropy decays slowly in the far-from-equilibrium phase and may persist until the relaxation time. Published by the American Physical Society 2024

Brewer, Jasmine (ORCID:0000000230840663)

Geometric Interpretation of the Cluster Location Problem Part II: Application to the Pahala, Hawaii, Earthquake Sequence

In the companion “Theory” article, we presented a new framing of the seismic location problem in terms of differential geometry (Harris et al., 2025). From that viewpoint, we developed a “project and correct” approach for estimating the relative locations of earthquakes. Here, in this study, we use project and correct to estimate high-precision relative locations of events from an earthquake sequence beneath the town of Pahala, Hawaii, using high-precision correlation-derived picks. The sequence was active from 2020 through 2022 and produced many highly correlated signals at Hawaii Volcano Observatory (HVO) stations on the island of Hawaii. The data we inverted consisted of 2882 events with observations at 5 HVO stations. For comparison with the travel-time image, we also produced conventional hypocenter solutions using both the Bayesloc program (Myers et al., 2007, 2009) and a purpose-built double-difference code. There were obvious structural elements in the resulting image, the resolution of which we used to test the performance of the project and the correct algorithm. For the projection step, we first produced a 3D local basis using an singular value decomposition (SVD) of the 2882 groups of times. Projection of the travel-time vectors into this basis resulted in an image with structures similar to those produced by our conventional locators, but with distortion as predicted by theory. Removing the distortion requires an inverse operator generated from the metric tensor at the geometric centroid of the events. We compared two approaches to obtaining such an inverse operator. The first uses an estimate of the geographic centroid of the event cloud from the centroid of the travel-time data. The second approach uses the centroid of the conventionally produced locations. The first approach produces a corrected image very similar to the conventional results, but with a rotation. The corrected image produced using the conventionally derived centroid is a near-exact match to the conventional locations.

Dodge, Douglas A. [Lawrence Livermore National Lab

Collocation methods for nonlinear differential equations on low-rank manifolds

We introduce new methods for integrating nonlinear differential equations on low-rank manifolds. These methods rely on interpolatory projections onto the tangent space, enabling low-rank time integration of vector fields that can be evaluated entry-wise. A key advantage of our approach is that it does not require the vector field to exhibit low-rank structure, thereby overcoming significant limitations of traditional dynamical low-rank methods based on orthogonal projection. To construct the interpolatory projectors, we develop a sparse tensor sampling algorithm based on the discrete empirical interpolation method (DEIM) that parameterizes tensor train manifolds and their tangent spaces with cross interpolation. Using these projectors, we propose two time integration schemes on low-rank tensor train manifolds. The first scheme integrates the solution at selected interpolation indices and constructs the solution with cross interpolation. The second scheme generalizes the well-known orthogonal projector-splitting integrator to interpolatory projectors. We demonstrate the proposed methods with applications to several tensor differential equations arising from the discretization of partial differential equations.

97 MATHEMATICS AND COMPUTING

Explicit simulation of the Brownian rotation of arbitrary shaped aerosol particles using quaternions

The shape of an aerosol particle strongly influences its mass and momentum transfer cross-sections, charging properties, and other physical properties. Here, we present an explicit time-stepping procedure to simulate the rotational Brownian motion of arbitrary shaped aerosol particles by solving Euler’s equation of rotation. A Langevin formulation of the rotation equations is used, wherein Brownian motion due to thermal collisions between a particle and background gas molecules is represented using a stochastic fluctuating torque and fluid resistance is included as a drag torque. To avoid singularities associated with describing the orientation of a shape with Euler angles, we employ a quaternion formulation that leads to first-order stochastic differential equations to describe the evolution of the angular position and angular velocity of a rigid body. We perform all the rotational dynamics calculations in the body-fixed frame of reference attached to the rotating shape whose basis vectors are the normalized eigenvectors of the inertia tensor of the particle. Numerical solutions to rotation under torque-free conditions, damped rotation without Brownian motion, and stochastic rotation for arbitrary shapes are presented and discussed. The presented method enables time-resolved simulation of Brownian rotation for direct comparison with experimentally measured trajectories or statistical measures. The second order accuracy of the used time-stepping procedure places a severe restriction on the timestep that can be used for obtaining accurate results. Animations of presented simulations are included for visualizing rotational motion at various gas pressures. To aid implementation, MATLAB ® codes are also provided. Extension to include translation Brownian motion is straightforward.

Roy, Mrittika

Observation of an Axial-Vector State in the Study of the Decay ψ ( 3686 ) → ϕ η η ′

Using ( 2712.4 ± 14.3 ) × 10 6 ψ ( 3686 ) events collected with the BESIII detector at BEPCII, a partial wave analysis of the decay ψ ( 3686 ) → ϕ η η ′ is performed with the covariant tensor approach. In addition to the established states h 1 ( 1900 ) and ϕ ( 2170 ) , an axial-vector state with a mass near 2.3 GeV / c 2 is observed for the first time. Its mass and width are measured to be 2316 ± 9 stat ± 3 0 syst MeV / c 2 and 89 ± 1 5 stat ± 2 6 syst MeV , respectively. The product branching fractions of B [ ψ ( 3686 ) → X ( 2300 ) η ′ ] B [ X ( 2300 ) → ϕ η ] and B [ ψ ( 3686 ) → X ( 2300 ) η ] B [ X ( 2300 ) → ϕ η ′ ] are determined to be ( 4.8 ± 1.3 stat ± 0.7 syst ) × 10 − 6 and ( 2.2 ± 0.7 stat ± 0.7 syst ) × 10 − 6 , respectively. The branching fraction B [ ψ ( 3686 ) → ϕ η η ′ ] is measured for the first time to be ( 3.14 ± 0.1 7 stat ± 0.2 4 syst ) × 10 − 5 . The first uncertainties are statistical and the second are systematic. Published by the American Physical Society 2025

Ablikim, M.

Impossibility of obtaining time-independent, three-dimensional, spherically symmetric densities of confined systems of relativistically moving constituents

The quantum-mechanical definition of probability, the uncertainty principle, and Poincaré invariance provide strong basic restrictions on the ability to define spatial densities associated with form factors describing the properties of confined systems of relativistically moving constituents. Despite this, many papers ignore one or more of these restrictions. Here I show how to obtain time-independent , two-dimensional densities that are consistent with the stated restrictions. This is done using the light-front, infinite momentum frame formalism. Two-dimensional density interpretations of the axial-vector form factor and all three gravitational form factors is obtained. The resulting mass radius is smaller than the charge radius. Additionally, an expression of a two-dimensional mass density related to the trace of the energy momentum tensor is obtained. I also show that all known methods for finding three-dimensional densities—using the Breit frame, Abel transformations, Wigner distributions, and spherically symmetric wave packets with vanishing spatial extent—violate the basic restrictions in different ways. Furthermore, the use of the latter leads to densities that vanish almost everywhere in space as time increases from an initial value.

form factors

Thermal Radiation Transport with Tensor Trains

We present a novel tensor network algorithm to solve the time-dependent, gray thermal radiation transport equation. The method invokes a tensor train (TT) decomposition for the specific intensity. The efficiency of this approach is dictated by the rank of the decomposition. When the solution is “low rank,” the memory footprint of the specific intensity solution vector may be significantly compressed. The algorithm, following a step-then-truncate approach of a traditional discrete ordinates method, operates directly on the compressed state vector, thereby enabling large speedups for low-rank solutions. To achieve these speedups, we rely on a recently developed rounding approach based on the Gram-SVD. We detail how familiar S N algorithms for (gray) thermal transport can be mapped to this TT framework and present several numerical examples testing both the optically thick and thin regimes. The TT framework finds low-rank structure and supplies up to ≃60× speedups and ≃1000× compressions for problems demanding large angle counts, thereby enabling previously intractable SN calculations and supplying a promising avenue to mitigate ray effects.

79 ASTRONOMY AND ASTROPHYSICS

Improving the Performance of NEML2 with Modern Graph Compilation Backends

NEML2 vectorizes constitutive-model evaluation for large-scale multiphysics simulation, using PyTorch as its tensor backend so that a batch of material-point updates runs on CPU or GPU through a single implementation. In the two prior reports in this series it was a C++-native library, deployed through TorchScript tracing and just-in-time (JIT) compilation; it has since been rewritten from the ground up into a Python-native library deployed through Ahead-of-Time Inductor (AOTInductor), a modern PyTorch graph-compilation backend. The rewrite is driven by a persistent tension, not a language preference: NEML2 composes constitutive models at runtime from a registry of small, independently-authored pieces, and that flexibility is difficult to reconcile with the compile-time knowledge an efficient GPU kernel needs. This report documents the rewrite and the investment that accompanied it: the AOTInductor export pipeline that turns a Python-authored model into a portable, Python-free compiled artifact loadable from pure C++; the eager and compiled runtimes and the new implicit solver layer built on them; a head-to-head benchmark of legacy JIT against AOTInductor; the physics-model catalog and its worked examples; the developer tooling; and the corresponding overhaul of MOOSE’s NEML2 integration that lets MOOSE consume it. A central objective is to examine whether modern PyTorch graph-compilation backends are effective for MOOSE GPU integration. The benchmark answers directly: AOTInductor outperforms legacy JIT on every GPU scenario measured, by 1.0–4.5×. Modern graph-compilation backends are effective for MOOSE GPU integration, and AOTInductor specifically – not compilation in the abstract – is why.

Hu, Gary (Tianchen) [Argonne National Laboratory (

Adiabatic fast passage spin manipulation measurements in solid polarized targets

Adiabatic fast passage (AFP) is a rapid method for reversing nuclear polarization and manipulating spin populations in polarized solid targets, avoiding the long repolarization times associated with dynamic nuclear polarization (DNP). We report AFP measurements in a 5 T, 1 K polarized-target system for irradiated 15 NH 3 , irradiated 14 ND 3 , and butanol-based materials prepared either with TEMPO doping or by irradiation. We also present a joint manipulated-lineshape analysis for spin-1 targets and demonstrate that vector and tensor polarizations can be extracted from AFP-manipulated deuteron NMR spectra even when the populations are not described by a single Boltzmann spin temperature. In conclusion, we report a reproducible polarization- and direction-dependent AFP response in a large irradiated 15 NH 3 sample. These ammonia results are presented as empirical observations under the specific sample–coil conditions of the experiment, with possible circuit-mediated mechanisms such as radiation damping or superradiant behavior discussed but not assigned as a definitive cause.

Adiabatic fast passage

Scalable learning of potentials to predict time-dependent Hartree–Fock dynamics

We propose a framework to learn the time-dependent Hartree–Fock (TDHF) inter-electronic potential of a molecule from its electron density dynamics. Although the entire TDHF Hamiltonian, including the inter-electronic potential, can be computed from first principles, we use this problem as a testbed to develop strategies that can be applied to learn a priori unknown terms that arise in other methods/approaches to quantum dynamics, e.g., emerging problems such as learning exchange–correlation potentials for time-dependent density functional theory. We develop, train, and test three models of the TDHF inter-electronic potential, each parameterized by a four-index tensor of size up to 60 × 60 × 60 × 60. Two of the models preserve Hermitian symmetry, while one model preserves an eight-fold permutation symmetry that implies Hermitian symmetry. Across seven different molecular systems, we find that accounting for the deeper eight-fold symmetry leads to the best-performing model across three metrics: training efficiency, test set predictive power, and direct comparison of true and learned inter-electronic potentials. All three models, when trained on ensembles of field-free trajectories, generate accurate electron dynamics predictions even in a field-on regime that lies outside the training set. To enable our models to scale to large molecular systems, we derive expressions for Jacobian-vector products that enable iterative, matrix-free training.

97 MATHEMATICS AND COMPUTING

Beyond-classical computation in quantum simulation

Quantum computers hold the promise of solving certain problems that lie beyond the reach of conventional computers. However, establishing this capability, especially for impactful and meaningful problems, remains a central challenge. Here, we show that superconducting quantum annealing processors can rapidly generate samples in close agreement with solutions of the Schrödinger equation. We demonstrate area-law scaling of entanglement in the model quench dynamics of two-, three-, and infinite-dimensional spin glasses, supporting the observed stretched-exponential scaling of effort for matrix-product-state approaches. We show that several leading approximate methods based on tensor networks and neural networks cannot achieve the same accuracy as the quantum annealer within a reasonable time frame. Thus, quantum annealers can answer questions of practical importance that may remain out of reach for classical computation.

King, Andrew D. [D-Wave Quantum Inc., Burnaby, BC

Numerically exact configuration interaction at quadrillion-determinant scale

The combinatorial growth of configuration interaction (CI) has long limited this formally exact quantum chemistry method to only the smallest molecules. Here, we report a numerically exact CI calculation exceeding one quadrillion (10 15 ) determinants, made possible by a lossless categorical compression strategy within the small-tensor-product distributed active space (STP-DAS) framework. This approach overcomes the traditional memory bottlenecks of CI by a numerically exact compression of the wavefunction representation and reformulating the most computationally demanding matrix–vector operations. Using this method, we performed a fully relativistic CI calculation of the ground state of HBrTe with over 10 15 complex-valued determinants in just 34.5 h on 1000 computing nodes—the largest CI calculation ever reported. We further achieved fast computation for systems with hundreds of billions of determinants on only a few compute nodes. Extensive benchmarks confirm that the method retains full numerical exactness while cutting memory and computational cost by orders of magnitude. Compared to previous state-of-the-art CI calculations, this work achieves a 1000 times increase in CI space, a 10 6 -fold increase in floating-point operations performed, and a 10 6 -fold improvement in computational speed.

Computational chemistry

Accelerating GNNs on GPU Sparse Tensor Cores through N:M Sparsity-Oriented Graph Reordering

Recent advancements in GPU hardware support have introduced the capability to leverage N:M sparse patterns for substantial performance gains. Graphs in Graph Neural Networks (GNNs) are typically sparse, but the sparsity is often irregular, not conforming to such sparse patterns. In this paper, we propose a novel graph reordering algorithm, the first of its kind, to reshape irregular graph data into the N:M structured sparse pattern at the tile level, allowing linear-algebra-based graph operations in GNNs to benefit from the N:M sparse hardware. The optimization is lossless, maintaining the accuracy of GNN. It can remove 98-100\% violations of the N:M sparse patterns at the vector level, and increase the proportion of conforming graphs in SuiteSparse collection from 5-9\% to 88.7-93.5\%. On A100 GPUs, the optimization accelerates Sparse Matrix Matrix (SpMM) by up to 43X (2.3X -- 7.5X on average) and speeds up the key graph operations in GNNs on real graphs by as much as 8.6X (3.5X on average).

artificial intelligence, graph neural networks

Scalable Quantum Monte Carlo Method for Polariton Chemistry via Mixed Block Sparsity and Tensor Hypercontraction Method

We present a reduced-scaling auxiliary-field quantum Monte Carlo (AFQMC) framework designed for large molecular systems and ensembles, with or without coupling to optical cavities. Our approach leverages the natural block sparsity of the Cholesky decomposition (CD) of electron repulsion integrals in molecular ensembles and employs tensor hypercontraction (THC) to efficiently compress low-rank Cholesky blocks. By representing the Cholesky vectors in a mixed format, keeping high-rank blocks in block-sparse form and compressing low-rank blocks with THC, we reduce the scaling of exchange-energy evaluation from quartic to robust cubic in the number of molecular orbitals N, while lowering memory from cubic toward quadratic. Benchmark analyses on one-, two-, and three-dimensional molecular ensembles (up to ∼1,200 orbitals) show that (a) the number of nonzeros in Cholesky tensors grows linearly with system size across dimensions; (b) the average numerical rank increases sublinearly and does not saturate at these sizes; and (c) rank heterogeneity─some blocks nearly full rank and many low rank, naturally motivates the proposed mixed block sparsity and THC scheme for efficient calculation of exchange energy. In conclusion, we demonstrate that the mixed scheme yields cubic wall-time scaling with favorable prefactors and preserves AFQMC accuracy.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH