Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “tensor cores”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Atypical phase-change alloy Ga 2 Te 3 : atomic structure, incipient nanotectonic nuclei, and multilevel writing

Emerging brain-inspired computing, including artificial optical synapses, photonic tensor cores, neuromorphic networks, etc., needs phase-change materials (PCMs) of the next generation with lower energy consumption and a wider temperature range for reliable long-term operation. Gallium tellurides with higher melting and crystallization temperatures appear to be promising candidates and enable achieving the necessary requirements. Here, using high energy X-ray diffraction and Raman spectroscopy supported by first-principles simulations, we show that vitreous g-Ga 2 Te 3 films essentially have a tetrahedral local structure and sp 3 hybridization, similar to those in the stable fcc Ga 2 Te 3 polymorph and in contrast to a vast majority of typical PCMs. Nevertheless, optical pump–probe laser experiments revealed high-contrast, fast and reversible multilevel SET-RESET transitions raising a question related to the phase change mechanism. A recently observed nanotectonic compression in bulk glassy Ga–Te alloys seems to be responsible for the PCM performance. Incipient nanotectonic nuclei, reminiscent of monoclinic high-pressure HP-Te II and rhombohedral HP-Ga 2 Te 3 , are present as minorities (2–4%) in g-Ga 2 Te 3 but are suggested to grow dramatically with increasing temperature while interacting with appropriate laser pulses. This leads to co-crystallization of HP-polymorphs amplified by a high internal local pressure reaching 4–8 GPa. The metallic HP-forms provide an increasing optical and electrical contrast, favorable for reliable PCM operations, and higher energy efficiency.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Interpretable AI forecasting for numerical relativity waveforms of quasicircular, spinning, nonprecessing binary black hole mergers

We present a deep-learning artificial intelligence model (AI) that is capable of learning and forecasting the late-inspiral, merger and ringdown of numerical relativity waveforms that describe quasicircular, spinning, nonprecessing binary black hole mergers. We used the NRHybSur3dq8 surrogate model to produce train, validation and test sets of ℓ = |m| = 2 waveforms that cover the parameter space of binary black hole mergers with mass ratios q ≤ 8 and individual spins |s$^{z}_ {(1,2)}$| ≤ 0.8. These waveforms cover the time range t ∊ [-5000 M, 130 M], where t = 0M marks the merger event, defined as the maximum value of the waveform amplitude. We harnessed the ThetaGPU supercomputer at the Argonne Leadership Computing Facility to train our AI model using a training set of 1.5 million waveforms. We used 16 NVIDIA DGX A100 nodes, each consisting of 8 NVIDIA A100 Tensor Core GPUs and 2 AMD Rome CPUs, to fully train our model within 3.5 h. Our findings show that artificial intelligence can accurately forecast the dynamical evolution of numerical relativity waveforms in the time range t ∊ [-100 M, 130 M]. Sampling a test set of 190,000 waveforms, we find that the average overlap between target and predicted waveforms is ≳99% over the entire parameter space under consideration. We also combined scientific visualization and accelerated computing to identify what components of our model take in knowledge from the early and late-time waveform evolution to accurately forecast the latter part of numerical relativity waveforms. This work aims to accelerate the creation of scalable, computationally efficient and interpretable artificial intelligence models for gravitational wave astrophysics

79 ASTRONOMY AND ASTROPHYSICS↗

Understanding Mixed Precision GEMM with MPGemmFI: Insights into Fault Resilience

Emerging deep learning workloads urgently need fast general matrix multiplication (GEMM). Thus, one of the critical features of machine-learning-specific accelerators such as NVIDIA Tensor Cores, AMD Matrix Cores, and Google TPUs is the support of mixed-precision enabled GEMM. For DNN models, lower-precision FP data formats and computation offer acceptable correctness but significant performance, area, and memory footprint improvement. While promising, the mixed-precision computation on error resilience remains unexplored. To this end, we develop a fault injection framework that systematically injects fault into the mixed-precision computation results. We investigate how the faults affect the accuracy of machine learning applications. Based on the characteristics of error resilience, we offer lightweight error detection and correction solutions that significantly improve the overall model accuracy by 75% if the models experience hardware faults. The solutions can be efficiently integrated into the accelerator's pipelines.

Fang, Bo↗

Bridging the Gap Between LLMs and LNS with Dynamic Data Format and Architecture Codesign

Deep neural networks (DNNs) have achieved tremendous success in the past few years. However, their training and inference demand exceptional computational and memory resources. Quantization has been shown as an effective approach to mitigate the cost, with the mainstream data types reduced from FP32 to FP16/BF16 and recently FP8 in the latest NVIDIA H100 GPUs. With increasingly aggressive quantization, however, the conventional floating-point formats suffer from limited precision in representing numbers around zero. Recently, NVIDIA demonstrated the potential of using a Logarithmic Number System (LNS) for the next generation of tensor cores. While LNS mitigates the hurdles in representing small numbers, in this work we observed a mismatch between LNS and the emerging Large Language Models (LLM), where LLM exhibits significant outliers when directly adopting the LNS format. In this paper, we present a data-format/architecture codesign to bright this gap. On the format side, we propose a dynamic LNS format to flexibly represent outliers at a higher precision, by exploiting asymmetry in the LNS representation and identifying outliers through a per-vector basis. On the architecture side, for demonstration, we realize the dynamic LNS format in a systolic array, which can handle the irregularity of the outliers at runtime. We implement our approach on an Alveo U280 FPGA as a prototype. Experimental results show that our design can effectively handle the outliers and resolve the mismatch between LNS and LLM, contributing to an accuracy improvement of 15.4% and 16% over the floating-point and the original LNS baselines, using four state-of-the-art LLM models. Our observation and design lay a solid foundation for the large-scale adoption of the LNS format in the next-generation deep learning hardware.

Haghi, Pouya↗

Real-time High-resolution X-Ray Computed Tomography

Computed Tomography (CT) serves as a key imaging technology that relies on computationally intensive filtering and back-projection algorithms for 3D image reconstruction. While conventional high-resolution image reconstruction (> 2K3) solutions provide quick results, they typically treat reconstruction as an offline workload to be performed remotely on large-scale HPC systems. The growing demand for post-construction AI-driven analytics and the need for real-time adjustments call for high-resolution reconstruction solutions that are feasible on local computing resources, i.e. a multi-GPU server at most. In this paper, we propose a novel approach that utilizes Tensor Cores to optimize image reconstruction without sacrificing precision. We also introduce a framework designed to enable real-time execution of end-to-end distributed image reconstruction in a multi-GPU environment. Evaluations conducted on a single Nvidia A100 and H100 GPU show performance improvements of 1.91 × and 2.15 × compared to highly optimized production libraries. Furthermore, our framework, when deployed on 8-card Nvidia A100 GPU system, demonstrates the ability to reconstruct real-world datasets into 20483 volumes (32 GB) in slightly more than one minute and 40963 volumes (256 GB) in 7 minutes.

Wu, Du↗

FTTN: Feature-Targeted Testing for Numerical Properties of NVIDIA & AMD Matrix Accelerators

FTTN is a test suite to evaluate the numerical behaviors of matrix accelerators of GPUs (NVIDIA Tensor Cores and AMD Matrix Cores) in a quick and simple setting. Matrix accelerators are heavily used in today's computationally intense applications to speed up matrix multiplications. This test suite provides a comprehensive study on the numerical behaviors of these accelerators, including support for subnormals, rounding modes, extra precision bits and FMA features. Is there

Laguna Peralta, Ignacio↗

Inference-Optimized AI and High Performance Computing for Gravitational Wave Detection at Scale

We introduce an ensemble of artificial intelligence models for gravitational wave detection that we trained in the Summit supercomputer using 32 nodes, equivalent to 192 NVIDIA V100 GPUs, within 2 h. Once fully trained, we optimized these models for accelerated inference using NVIDIA TensorRT. We deployed our inference-optimized AI ensemble in the ThetaGPU supercomputer at Argonne Leadership Computer Facility to conduct distributed inference. Using the entire ThetaGPU supercomputer, consisting of 20 nodes each of which has 8 NVIDIA A100 Tensor Core GPUs and 2 AMD Rome CPUs, our NVIDIA TensorRT-optimized AI ensemble processed an entire month of advanced LIGO data (including Hanford and Livingston data streams) within 50 s. Our inference-optimized AI ensemble retains the same sensitivity of traditional AI models, namely, it identifies all known binary black hole mergers previously identified in this advanced LIGO dataset and reports no misclassifications, while also providing a 3X inference speedup compared to traditional artificial intelligence models. We used time slides to quantify the performance of our AI ensemble to process up to 5 years worth of advanced LIGO data. In this synthetically enhanced dataset, our AI ensemble reports an average of one misclassification for every month of searched advanced LIGO data. We also present the receiver operating characteristic curve of our AI ensemble using this 5 year long advanced LIGO dataset. This approach provides the required tools to conduct accelerated, AI-driven gravitational wave detection at scale.

97 MATHEMATICS AND COMPUTING↗

Design of a Low-Cost, Submersible, Digital Holographic Microscope for in Situ Microbial Imaging

The methodologies for studying marine microbiology typically consist of utilizing instrumentation within a laboratory. This typically requires extracting a sample from its place of origin prior to examination, which may be days after collection. Oftentimes, the solution is to bring the lab to the ocean, which may be costly and provide further limitations for a sterile and stable laboratory environment. Here we present a low-cost, submersible, digital holographic microscope (DHM) designed to image marine microorganisms (such as bacteria and plankton) in their natural underwater environment. Our instrument eliminates the need to transport samples and allows for instantaneous data collection of microbes in-situ. The DHM achieves sub-micron spatial resolution and is paired with artificial intelligence for the detection and tracking of specimens to reduce the overall collected data. This instrument also aims to reduce the cost of manufacturing and field expenses relative to marine microbiological research. “Off the shelf” components were selected in the design process of this instrument which allows us to achieve precise results without sacrificing data quality. The DHM itself costs under one thousand dollars and features a low-cost high-resolution camera, the Arducam MT9J001. Included in the design were five main subsystems: optical, mechanical, electrical, power, and machine learning. Our on board computer and artificial intelligence consist of a Raspberry Pi 4 (8 Gb) and Google Coral USB Tensorflow accelerator. Instrument testing has successfully proven our abilities of data acquisition for at least two hours in depths of at least forty meters below sea level. Additionally, our artificial intelligence system is currently capable of tracking up to ten areas of interest in a fraction of a second with over ninety percent confidence via the neural net driven by the tensor cores on the Google Coral. Furthermore, we have demonstrated the versatility of our instrument by mounting it on an ocean-going ROV, the BlueRov2 by BlueRobotics. Our tests on the BlueRov2 exemplified the cost-effective nature of a submersible and reusable instrument that can be implemented in moderate environments and on most vessels.

Wallace, James Kent↗

Kinematics and dynamics of disclination lines in three-dimensional nematics

An exact kinematic law for the motion of disclination lines in nematic liquid crystals as a function of the tensor order parameter Q is derived. Unlike other order parameter fields that become singular at their respective defect cores, the tensor order parameter remains regular. Following earlier experimental and theoretical work, the disclination core is defined to be the line where the uniaxial and biaxial order parameters are equal, or equivalently, where the two largest eigenvalues of Q cross. This allows an exact expression relating the velocity of the line to spatial and temporal derivatives of Q on the line, to be specified by a dynamical model for the evolution of the nematic. By introducing a linear core approximation for Q, analytical results are given for several prototypical configurations, including line interactions and motion, loop annihilation, and the response to external fields and shear flows. Behaviour that follows from topological constraints or defect geometry is highlighted. Finally, the analytic results are shown to be in agreement with three-dimensional numerical calculations based on a singular Maier–Saupe free energy that allows for anisotropic elasticity.

36 MATERIALS SCIENCE↗

Entanglement structures in quantum field theories: Negativity cores and bound entanglement in the vacuum

Here, the many-body entanglement between two finite (size-d) disjoint vacuum regions of noninteracting lattice scalar field theory in one spatial dimension, i.e., a (d A × d B ) mixed Gaussian continuous variable system, is locally transformed into a tensor-product core of (1 A × 1 B ) mixed entangled pairs. Accessible entanglement within these core pairs exhibits an exponential hierarchy and as such identifies the structure of dominant region modes from which vacuum entanglement could be extracted into a spatially separated pair of quantum detectors. Beyond the core, the remaining modes of the halo are determined to be AB separable in isolation, as well as separable from the core. However, state preparation protocols that distribute entanglement in the form of (1 A × 1 B ) mixed core pairs are found to require additional entanglement in the halo that is obscured by classical correlations. This inaccessible (bound) halo entanglement is found to mirror the accessible entanglement, but with a step behavior as the continuum is approached. It remains possible that alternate initialization protocols that do not utilize the exponential hierarchy of core-pair entanglement may require less inaccessible entanglement. Entanglement consolidation is expected to persist in higher dimensions and may aid classical and quantum simulations of asymptotically free gauge field theories, such as quantum chromodynamics.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

An Incremental Tensor Train Decomposition Algorithm

We present a new algorithm for incrementally updating the tensor train decomposition of a stream of tensor data. This new algorithm, called the tensor train incremental core expansion (TT-ICE) improves upon the current state-of-the-art algorithms for compressing in tensor train format by developing a new adaptive approach that incurs significantly slower rank growth and guarantees compression accuracy. This capability is achieved by limiting the number of new vectors appended to the TT-cores of an existing accumulation tensor after each data increment. These vectors represent directions orthogonal to the span of existing cores and are limited to those needed to represent a newly arrived tensor to a target accuracy. We provide two versions of the algorithm: TT-ICE and TT-ICE accelerated with heuristics (TT-ICE*). Here, we provide a proof of correctness for TT-ICE and empirically demonstrate the performance of the algorithms in compressing large-scale video and scientific simulation datasets. Compared to existing approaches that also use rank adaptation, TT-ICE* achieves 57× higher compression and up to 95% reduction in computational time.

97 MATHEMATICS AND COMPUTING↗

Three-dimensional atomic structure and local chemical order of medium- and high-entropy nanoalloys

Medium- and high-entropy alloys (M/HEAs) mix several principal elements with near-equiatomic composition and represent a model-shift strategy for designing previously unknown materials in metallurgy, catalysis and other fields. One of the core hypotheses of M/HEAs is lattice distortion, which has been investigated by different numerical and experimental techniques. However, determining the three-dimensional (3D) lattice distortion in M/HEAs remains a challenge. Moreover, the presumed random elemental mixing in M/HEAs has been questioned by X-ray and neutron studies, atomistic simulations, energy dispersive spectroscopy and electron diffraction, which suggest the existence of local chemical order in M/HEAs. However, direct experimental observation of the 3D local chemical order has been difficult because energy dispersive spectroscopy integrates the composition of atomic columns along the zone axes and diffuse electron reflections may originate from planar defects instead of local chemical order. Here, in this work, we determine the 3D atomic positions of M/HEA nanoparticles using atomic electron tomography and quantitatively characterize the local lattice distortion, strain tensor, twin boundaries, dislocation cores and chemical short-range order (CSRO). We find that the high-entropy alloys have larger local lattice distortion and more heterogeneous strain than the medium-entropy alloys and that strain is correlated to CSRO. We also observe CSRO-mediated twinning in the medium-entropy alloys, that is, twinning occurs in energetically unfavoured CSRO regions but not in energetically favoured CSRO ones, which represents, to our knowledge, the first experimental observation of correlating local chemical order with structural defects in any material. We expect that this work will not only expand our fundamental understanding of this important class of materials but also provide the foundation for tailoring M/HEA properties through engineering lattice distortion and local chemical order.

36 MATERIALS SCIENCE↗

Nonnegative canonical tensor decomposition with linear constraints: nnCANDELINC

Abstract There is an emerging interest for tensor factorization applications in big‐data analytics and machine learning. To speed up the factorization of extra‐large datasets, organized in multidimensional arrays (also known as tensors), easy to compute compression‐based tensor representations, such as, Tucker and tensor train formats, are used to approximate the initial large‐tensor. Further, tensor factorization is used to extract latent features that can facilitate discoveries of new mechanisms and signatures hidden in the data, where the explainability of the latent features is of principal importance. Nonnegative tensor factorization extracts latent features that are naturally sparse and parts of the data, which makes them easily interpretable. However, to take into account available domain knowledge and subject matter expertise, often additional constraints need to be imposed, which lead us to canonical decomposition with linear constraints (CANDELINC), a canonical polyadic decomposition with rank deficient factors. In CANDELINC, Tucker compression is used as a preprocessing step, which lead to a larger residual error but to more explainable latent features. Here, we propose a nonnegative CANDELINC (nnCANDELINC) accomplished via a specific nonnegative Tucker decomposition; we refer to as minimal or canonical nonnegative Tucker. We derive several results required to understand the specificity of nnCANDELINC, focusing on the difficulties of preserving the nonnegative rank of a tensor to its Tucker core and comparing the real valued to nonnegative case. Finally, we demonstrate nnCANDELINC performance on synthetic and real‐world examples.

97 MATHEMATICS AND COMPUTING↗

Searching for three-nucleon short-range correlations

Electron scattering measurements from high-momentum nucleons in nuclei at SLAC and Jefferson Lab (JLab) have shown that these nucleons are generally associated with two-nucleon short-range correlations (2N-SRCs). These SRCs are formed when two nucleons in the nucleus interact at short distance via the strong tensor attraction or repulsive core of the NN potential. A series of measurements at JLab have mapped out the A dependence and isospin dependence of 2N-SRCs, and have begun to map out their momentum structure. However, we do not yet know if 3N-SRCs, similar high-momentum configurations of three nucleons, play an important role in nuclei. Here, we summarize here previous attempts to isolate 3N-SRCs, go over the limitations of these previous attempts, and discuss the present and near-term prospects for searching for 3N-SRCs, mapping out their A dependence in nuclei, and constraining their isospin and momentum structure.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

TensorID v1.0

This Python software package includes new and efficient algorithms for satellite and core interpolative decomposition of tensor data. In general, these algorithms target high-dimensional data reduction and compression. The software is purely numerical and can be applied by others to many important sources of tensor data generated by computation or experiment.

Zhang, Yifan [Lawrence Berkeley National Laborator↗

QSpace - An open-source tensor library for Abelian and non-Abelian symmetries

This is the documentation for the tensor library QSpace (v4.0), a toolbox to exploit ‘quan tum symmetry spaces’ in tensor network states in the quantum many-body context. QSpace permits arbitrary combinations of symmetries including the abelian symmetries $\mathbb{Z}_n$ and U(1), as well as all non-abelian symmetries based on the semisimple classical Lie algebras: A n , B n , C n , and D n , or respectively, the special unitary group SU(n), the odd orthogonal group SO(2n+1), the symplectic group Sp(2n), and the even orthogonal group SO(2n). The code (C++ embedded via the MEX interface into Matlab) is available open source as of QSpace v4.0 on bitbucket under the Apache 2.0 license. QSpace is designed as a bottom-up approach for non-abelian symmetries. It starts from the defining representation and the respective Lie algebra. By explicitly comput ing and tabulating generalized Clebsch-Gordan coefficient tensors, QSpace is versatile in the type of operations that it can perform across all symmetries. At the level of an ap plication, much of the symmetry-related details are hidden within the QSpace C++ core libraries. Hence when developing tensor network algorithms with QSpace, these can be coded (nearly) as if there are no symmetries at all, despite being able to fully exploit general non-abelian symmetries.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Measurements of the quantum geometric tensor in solids

Understanding the geometric properties of quantum states and their implications in fundamental physical phenomena is a core aspect of contemporary physics. The quantum geometric tensor (QGT) is a central physical object in this regard, encoding complete information about the geometry of the quantum state. The imaginary part of the QGT is the well-known Berry curvature, which plays an integral role in the topological magnetoelectric and optoelectronic phenomena. The real part of the QGT is the quantum metric, whose importance has come to prominence recently, giving rise to a new set of quantum geometric phenomena such as anomalous Landau levels, flat band superfluidity, excitonic Lamb shifts and nonlinear Hall effect. Despite the central importance of the QGT, its experimental measurements have been restricted only to artificial two-level systems. Here, in this work, we develop a framework to measure the QGT in crystalline solids using polarization-, spin- and angle-resolved photoemission spectroscopy. Using this framework, we demonstrate the effective reconstruction of the QGT in the kagome metal CoSn, which hosts topological flat bands. Establishing this momentum- and energy-resolved spectroscopic probe of the QGT is poised to significantly advance our understanding of quantum geometric responses in a wide range of crystalline systems.

36 MATERIALS SCIENCE↗