Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “tensor cores”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Real-time High-resolution X-Ray Computed Tomography

Computed Tomography (CT) serves as a key imaging technology that relies on computationally intensive filtering and back-projection algorithms for 3D image reconstruction. While conventional high-resolution image reconstruction (> 2K3) solutions provide quick results, they typically treat reconstruction as an offline workload to be performed remotely on large-scale HPC systems. The growing demand for post-construction AI-driven analytics and the need for real-time adjustments call for high-resolution reconstruction solutions that are feasible on local computing resources, i.e. a multi-GPU server at most. In this paper, we propose a novel approach that utilizes Tensor Cores to optimize image reconstruction without sacrificing precision. We also introduce a framework designed to enable real-time execution of end-to-end distributed image reconstruction in a multi-GPU environment. Evaluations conducted on a single Nvidia A100 and H100 GPU show performance improvements of 1.91 × and 2.15 × compared to highly optimized production libraries. Furthermore, our framework, when deployed on 8-card Nvidia A100 GPU system, demonstrates the ability to reconstruct real-world datasets into 20483 volumes (32 GB) in slightly more than one minute and 40963 volumes (256 GB) in 7 minutes.

Wu, Du↗

FTTN: Feature-Targeted Testing for Numerical Properties of NVIDIA & AMD Matrix Accelerators

FTTN is a test suite to evaluate the numerical behaviors of matrix accelerators of GPUs (NVIDIA Tensor Cores and AMD Matrix Cores) in a quick and simple setting. Matrix accelerators are heavily used in today's computationally intense applications to speed up matrix multiplications. This test suite provides a comprehensive study on the numerical behaviors of these accelerators, including support for subnormals, rounding modes, extra precision bits and FMA features. Is there

Laguna Peralta, Ignacio↗

Kinematics and dynamics of disclination lines in three-dimensional nematics

An exact kinematic law for the motion of disclination lines in nematic liquid crystals as a function of the tensor order parameter Q is derived. Unlike other order parameter fields that become singular at their respective defect cores, the tensor order parameter remains regular. Following earlier experimental and theoretical work, the disclination core is defined to be the line where the uniaxial and biaxial order parameters are equal, or equivalently, where the two largest eigenvalues of Q cross. This allows an exact expression relating the velocity of the line to spatial and temporal derivatives of Q on the line, to be specified by a dynamical model for the evolution of the nematic. By introducing a linear core approximation for Q, analytical results are given for several prototypical configurations, including line interactions and motion, loop annihilation, and the response to external fields and shear flows. Behaviour that follows from topological constraints or defect geometry is highlighted. Finally, the analytic results are shown to be in agreement with three-dimensional numerical calculations based on a singular Maier–Saupe free energy that allows for anisotropic elasticity.

36 MATERIALS SCIENCE↗

Entanglement structures in quantum field theories: Negativity cores and bound entanglement in the vacuum

Here, the many-body entanglement between two finite (size-d) disjoint vacuum regions of noninteracting lattice scalar field theory in one spatial dimension, i.e., a (d A × d B ) mixed Gaussian continuous variable system, is locally transformed into a tensor-product core of (1 A × 1 B ) mixed entangled pairs. Accessible entanglement within these core pairs exhibits an exponential hierarchy and as such identifies the structure of dominant region modes from which vacuum entanglement could be extracted into a spatially separated pair of quantum detectors. Beyond the core, the remaining modes of the halo are determined to be AB separable in isolation, as well as separable from the core. However, state preparation protocols that distribute entanglement in the form of (1 A × 1 B ) mixed core pairs are found to require additional entanglement in the halo that is obscured by classical correlations. This inaccessible (bound) halo entanglement is found to mirror the accessible entanglement, but with a step behavior as the continuum is approached. It remains possible that alternate initialization protocols that do not utilize the exponential hierarchy of core-pair entanglement may require less inaccessible entanglement. Entanglement consolidation is expected to persist in higher dimensions and may aid classical and quantum simulations of asymptotically free gauge field theories, such as quantum chromodynamics.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

An Incremental Tensor Train Decomposition Algorithm

We present a new algorithm for incrementally updating the tensor train decomposition of a stream of tensor data. This new algorithm, called the tensor train incremental core expansion (TT-ICE) improves upon the current state-of-the-art algorithms for compressing in tensor train format by developing a new adaptive approach that incurs significantly slower rank growth and guarantees compression accuracy. This capability is achieved by limiting the number of new vectors appended to the TT-cores of an existing accumulation tensor after each data increment. These vectors represent directions orthogonal to the span of existing cores and are limited to those needed to represent a newly arrived tensor to a target accuracy. We provide two versions of the algorithm: TT-ICE and TT-ICE accelerated with heuristics (TT-ICE*). Here, we provide a proof of correctness for TT-ICE and empirically demonstrate the performance of the algorithms in compressing large-scale video and scientific simulation datasets. Compared to existing approaches that also use rank adaptation, TT-ICE* achieves 57× higher compression and up to 95% reduction in computational time.

97 MATHEMATICS AND COMPUTING↗

Three-dimensional atomic structure and local chemical order of medium- and high-entropy nanoalloys

Medium- and high-entropy alloys (M/HEAs) mix several principal elements with near-equiatomic composition and represent a model-shift strategy for designing previously unknown materials in metallurgy, catalysis and other fields. One of the core hypotheses of M/HEAs is lattice distortion, which has been investigated by different numerical and experimental techniques. However, determining the three-dimensional (3D) lattice distortion in M/HEAs remains a challenge. Moreover, the presumed random elemental mixing in M/HEAs has been questioned by X-ray and neutron studies, atomistic simulations, energy dispersive spectroscopy and electron diffraction, which suggest the existence of local chemical order in M/HEAs. However, direct experimental observation of the 3D local chemical order has been difficult because energy dispersive spectroscopy integrates the composition of atomic columns along the zone axes and diffuse electron reflections may originate from planar defects instead of local chemical order. Here, in this work, we determine the 3D atomic positions of M/HEA nanoparticles using atomic electron tomography and quantitatively characterize the local lattice distortion, strain tensor, twin boundaries, dislocation cores and chemical short-range order (CSRO). We find that the high-entropy alloys have larger local lattice distortion and more heterogeneous strain than the medium-entropy alloys and that strain is correlated to CSRO. We also observe CSRO-mediated twinning in the medium-entropy alloys, that is, twinning occurs in energetically unfavoured CSRO regions but not in energetically favoured CSRO ones, which represents, to our knowledge, the first experimental observation of correlating local chemical order with structural defects in any material. We expect that this work will not only expand our fundamental understanding of this important class of materials but also provide the foundation for tailoring M/HEA properties through engineering lattice distortion and local chemical order.

36 MATERIALS SCIENCE↗

Nonnegative canonical tensor decomposition with linear constraints: nnCANDELINC

Abstract There is an emerging interest for tensor factorization applications in big‐data analytics and machine learning. To speed up the factorization of extra‐large datasets, organized in multidimensional arrays (also known as tensors), easy to compute compression‐based tensor representations, such as, Tucker and tensor train formats, are used to approximate the initial large‐tensor. Further, tensor factorization is used to extract latent features that can facilitate discoveries of new mechanisms and signatures hidden in the data, where the explainability of the latent features is of principal importance. Nonnegative tensor factorization extracts latent features that are naturally sparse and parts of the data, which makes them easily interpretable. However, to take into account available domain knowledge and subject matter expertise, often additional constraints need to be imposed, which lead us to canonical decomposition with linear constraints (CANDELINC), a canonical polyadic decomposition with rank deficient factors. In CANDELINC, Tucker compression is used as a preprocessing step, which lead to a larger residual error but to more explainable latent features. Here, we propose a nonnegative CANDELINC (nnCANDELINC) accomplished via a specific nonnegative Tucker decomposition; we refer to as minimal or canonical nonnegative Tucker. We derive several results required to understand the specificity of nnCANDELINC, focusing on the difficulties of preserving the nonnegative rank of a tensor to its Tucker core and comparing the real valued to nonnegative case. Finally, we demonstrate nnCANDELINC performance on synthetic and real‐world examples.

97 MATHEMATICS AND COMPUTING↗

Searching for three-nucleon short-range correlations

Electron scattering measurements from high-momentum nucleons in nuclei at SLAC and Jefferson Lab (JLab) have shown that these nucleons are generally associated with two-nucleon short-range correlations (2N-SRCs). These SRCs are formed when two nucleons in the nucleus interact at short distance via the strong tensor attraction or repulsive core of the NN potential. A series of measurements at JLab have mapped out the A dependence and isospin dependence of 2N-SRCs, and have begun to map out their momentum structure. However, we do not yet know if 3N-SRCs, similar high-momentum configurations of three nucleons, play an important role in nuclei. Here, we summarize here previous attempts to isolate 3N-SRCs, go over the limitations of these previous attempts, and discuss the present and near-term prospects for searching for 3N-SRCs, mapping out their A dependence in nuclei, and constraining their isospin and momentum structure.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

TensorID v1.0

This Python software package includes new and efficient algorithms for satellite and core interpolative decomposition of tensor data. In general, these algorithms target high-dimensional data reduction and compression. The software is purely numerical and can be applied by others to many important sources of tensor data generated by computation or experiment.

Zhang, Yifan [Lawrence Berkeley National Laborator↗

QSpace - An open-source tensor library for Abelian and non-Abelian symmetries

This is the documentation for the tensor library QSpace (v4.0), a toolbox to exploit ‘quan tum symmetry spaces’ in tensor network states in the quantum many-body context. QSpace permits arbitrary combinations of symmetries including the abelian symmetries $\mathbb{Z}_n$ and U(1), as well as all non-abelian symmetries based on the semisimple classical Lie algebras: A n , B n , C n , and D n , or respectively, the special unitary group SU(n), the odd orthogonal group SO(2n+1), the symplectic group Sp(2n), and the even orthogonal group SO(2n). The code (C++ embedded via the MEX interface into Matlab) is available open source as of QSpace v4.0 on bitbucket under the Apache 2.0 license. QSpace is designed as a bottom-up approach for non-abelian symmetries. It starts from the defining representation and the respective Lie algebra. By explicitly comput ing and tabulating generalized Clebsch-Gordan coefficient tensors, QSpace is versatile in the type of operations that it can perform across all symmetries. At the level of an ap plication, much of the symmetry-related details are hidden within the QSpace C++ core libraries. Hence when developing tensor network algorithms with QSpace, these can be coded (nearly) as if there are no symmetries at all, despite being able to fully exploit general non-abelian symmetries.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Measurements of the quantum geometric tensor in solids

Understanding the geometric properties of quantum states and their implications in fundamental physical phenomena is a core aspect of contemporary physics. The quantum geometric tensor (QGT) is a central physical object in this regard, encoding complete information about the geometry of the quantum state. The imaginary part of the QGT is the well-known Berry curvature, which plays an integral role in the topological magnetoelectric and optoelectronic phenomena. The real part of the QGT is the quantum metric, whose importance has come to prominence recently, giving rise to a new set of quantum geometric phenomena such as anomalous Landau levels, flat band superfluidity, excitonic Lamb shifts and nonlinear Hall effect. Despite the central importance of the QGT, its experimental measurements have been restricted only to artificial two-level systems. Here, in this work, we develop a framework to measure the QGT in crystalline solids using polarization-, spin- and angle-resolved photoemission spectroscopy. Using this framework, we demonstrate the effective reconstruction of the QGT in the kagome metal CoSn, which hosts topological flat bands. Establishing this momentum- and energy-resolved spectroscopic probe of the QGT is poised to significantly advance our understanding of quantum geometric responses in a wide range of crystalline systems.

36 MATERIALS SCIENCE↗

SymProp: Scaling Sparse Symmetric Tucker Decomposition via Symmetry Propagation

Sparse symmetric tensors are an important class of tensors, and their decompositions serve as powerful tools for revealing low-rank structures. This paper introduces SymProp, a novel approach for scaling sparse symmetric Tucker decomposition by propagating symmetry through intermediate computations. SymProp optimizes two key computational kernels: Sparse Symmetric Tensor Times Same Matrix chain (S3 TTMc) for Higher-Order Orthogonal Iteration (HOOI) and Sparse Symmetric Tensor Times Same Matrix chain Times Core (S3 TTMcTC) for Higher-Order QR Iteration (HOQRI). Our method employs a metaprogramming-based index iteration approach to efficiently handle the upper triangular parts of intermediate dense symmetric tensors. SymProp achieves up to 50.9× speedup over SPLATT and up to 360.8× over Compressed Sparse Symmetric (CSS) format on the S3 TTMc operation. Moreover, our S3 TTMc and S3 TTMcTC implementations support tensor orders four levels higher than state-of-the-art methods. Our HOQRI demonstrates superior scalability and up to a 33.6× speedup over optimized HOOI. By enabling more scalable Tucker decompositions for higher orders, decomposition ranks, and dimension sizes, SymProp opens new possibilities for analyzing complex hypergraph structures in fields such as network science, data mining, and machine learning.

Li, Zecheng [North Carolina State University]↗

Quantifying the dislocation content of atomically resolved grain boundary line defects using the Nye tensor

The Nye tensor, which quantifies the density of Burgers vector for a given dislocation line direction, can be effectively used to characterize dislocation content in bulk crystals from atomic-resolution transmission electron microscopy images. The Nye tensor can be calculated from these images, in part because the reference state is simply defined by the lattice of the perfect crystal. The application of the Nye tensor to interfacial line defects, for which the natural reference state is the dichromatic pattern of the two grains in their reference orientation, poses additional challenges. In this work, we present a method that employs the Nye tensor to characterize the edge dislocation content of line defects at grain boundaries from atomic-resolution images. This approach enables us to rapidly characterize all edge dislocation content along a grain boundary. Additionally, the Nye tensor provides information about line defect core structure. Finally, we demonstrate this method on two exemplar defects: a twin boundary disconnection and a facet junction in face-centered cubic Au.

Crystallographic defects↗

Ion cyclotron heating at high plasma density in Proto-MPEX

The physics of ion cyclotron heating (ICH) relevant to the steady-state linear machine MPEX (Material Plasma Exposure eXperiment) has been explored in its predecessor, short-pulse device: Proto-MPEX. MPEX will utilize fundamental ICH to increase heat flux at the target and produce ion temperatures and velocity distributions with improved fidelity to those found in a tokamak divertor region, in comparison to those produced by substrate biasing. In the experiments on Proto-MPEX described here, bulk ion temperatures up to ~15 eV have been achieved with 20 kW net ICH power at 6.5 MHz, using ICH heating of a deuterium plasma produced by a helicon plasma source. The heat flux at the target has been observed to increase throughout the plasma cross section, including in the core region. Core $T_i$ and target heat flux are observed to scale linearly with injected ICH power. Further, measurements of plasma loading and target heat flux as a function of the magnetic field strength at the antenna, together with modeling of the wave propagation from the antenna to the ion cyclotron resonance using the ANTENA and COMSOL codes with a warm plasma dielectric tensor, indicate that power is coupled to the core plasma via fast wave excitation of a kinetic Alfvén wave.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

One- and two-dimensional higher-point conformal blocks as free-particle wavefunctions in $$ {\textrm{AdS}}_3^{\otimes m} $$

Abstract We establish that all of the one- and two-dimensional global conformal blocks are, up to some choice of prefactor, free-particle wavefunctions in tensor products of AdS 3 or limits thereof. Our first core observation is that the six-point comb-channel conformal blocks correspond to free-particle wavefunctions on an AdS 3 constructed directly in cross-ratio space. This construction generalizes to blocks for a special class of diagrams, which are determined as free-particle wavefunctions in tensor products of AdS 3 . Conformal blocks for all the remaining topologies are obtained as limits of the free wavefunctions mentioned above. Our results show directly that the integrable models associated with all one- and two-dimensional conformal blocks can be seen as limits of free theory, and manifest a relation between AdS and CFT kinematics that lies outside of the standard AdS/CFT dictionary. We complete the discussion by providing explicit Feynman-like rules that can be used to work out blocks for all topologies, as well as a Mathematica notebook that allows simple computation of Casimir equations and series expansions for blocks, by requiring just an OPE diagram as input.

Physics↗

Continued Verification of MOOSE Structural Mechanics Tools for Modeling Core Bowing Phenomena in Fast Reactors

Under the U.S. Department of Energy Office of Nuclear Energy’s Advanced Modeling and Simulation (NEAMS) Program, an integrated multiphysics approach is being developed to model the core bowing phenomena important to liquid metal-cooled fast reactors. Core bowing is an important passive safety mechanism whereby increased power (which leads to temperature and flux gradients) influences the core to bow into less reactive configurations when the restraint system is properly designed. The phenomenon includes a complex interplay of radiation transport, duct temperature calculations involving fluid flow and heat transfer, and thermo-mechanical responses to the induced temperature and flux gradients. Structural material properties are also important to determining inelastic response to longer term flux gradients which cause irradiation creep and swelling. While core bowing provides a strong negative reactivity feedback when the restraint system is designed properly, it also results in additional forces between assemblies which increase the loads required to extricate them during refueling or control rod movement. Therefore, the restraint system must be designed with these tradeoffs in mind. The first stage of the work, which commenced in FY21 and continues through FY22, assesses thermo-mechanical modeling tools for producing core bowing predictions consistent with conventional tools. The Multiphysics Object Oriented Simulation Environment (MOOSE) Tensor Mechanics and Contact Modules are employed. This status report describes work on additional thermo-mechanical benchmark verification problems with increased complexity from the examples demonstrated in FY21. Several benchmark verification examples were selected from the IAEA verification and validation report. These examples involve clusters of ducts representative of a sector of a hexagonal reactor core which bow into each other and cause contact and load pad elevations, as well as single ducts subjected to irradiation fields undergoing swelling and subsequent bowing. The MOOSE-based results were compared to both IAEA benchmark participants’ results, analytic equations as available, and NUBOW-3D, a beam model code developed by Argonne National Laboratory. In every case, the MOOSE results agreed with other simulations results, providing additional verification basis of the tools for this particular physics application.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Electrical conductivity of a warm neutron star crust in magnetic fields: Neutron-drip regime

We compute the anisotropic electrical conductivity tensor of the inner crust of a compact star at nonzero temperature by extending a previous work on the conductivity of the outer crust. The physical scenarios, where such crust is formed, involve protoneutron stars born in supernova explosions, binary neutron star mergers, and accreting neutron stars. The temperature-density range studied covers the transition from a semidegenerate to a highly degenerate electron gas and assumes that the nuclei form a liquid, i.e., the temperature is above the melting temperature of the lattice of nuclei. The electronic transition probabilities include (i) the screening of electron-ion interaction in the hard-thermal-loop approximation for the QED plasma, (ii) the correlations of the ionic component in a one-component plasma, and (iii) finite nuclear size effects. The conductivity tensor is obtained from the Boltzmann kinetic equation in relaxation time approximation accounting for the anisotropy introduced by a magnetic field. The sensitivity of the results towards the matter composition of the inner crust is explored by using several compositions of the inner crust, which were obtained using different nuclear interactions and methods of solving the many-body problem. The standard deviations of relaxation time and components of the conductivity tensor from the average are below ≤25% except close to crust-core transition, where nonspherical nuclear structures are expected. Finally, our results can be used in dissipative magnetohydrodynamics simulations of warm compact stars.

Physics↗