Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “tensor cores”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Atomic Structure and Dynamics of Unusual and Wide-Gap Phase-Change Chalcogenides: A GeTe 2 Case

Brain-inspired computing, reconfigurable optical metamaterials, photonic tensor cores, and many other advanced applications require next-generation phase-change materials (PCMs) with better energy efficiency and a wider thermal and spectral range for reliable operations. Germanium ditelluride (GeTe 2 ), with higher thermal stability and a larger bandgap compared to current benchmark PCMs, appears promising for THz metasurfaces and the controlled crystallization of atomically thin 2D materials. Using high-energy X-Ray diffraction supported by first-principles simulation, the atomic structure in semiconducting pulsed laser deposition films and metallic high-temperature liquids is investigated. The results suggest that the structural and chemical metastability of GeTe 2 , leading to disproportionation into GeTe and Te, is related to high internal pressure during a semiconductor–metal transition, presumably occurring in the supercooled melt. Similar phenomena are expected for canonical GeS 2 and GeSe 2 under high temperatures and pressures.

74 ATOMIC AND MOLECULAR PHYSICS↗

Fragile-to-strong transition in liquid As 2 S 3 under pressure: The effect of melt metallization

The well-known classification of glass-forming melts into fragile and strong liquids has several notable exceptions, including water, silica, and certain phase-change materials (PCMs). These exceptional fluid systems exhibit a fragile-to-strong transition (FST) upon cooling: a transformation from a high-temperature liquid with fast atomic dynamics, low viscosity, and low flow activation energy, to a viscous supercooled melt with high energy barriers near the glass transition temperature T g . This behavior is critically important for non-volatile memories, photonic tensor cores, reconfigurable metamaterials, and other devices, that use PCMs, enabling nanosecond-scale crystallization in the fragile regime and long data retention in the strong regime near or below T g . A significant structural transformation is expected between these two viscosity regimes, along with a semiconductor-metal (SC-M) transition upon heating, driven by high internal pressure and associated density increase. By applying high external pressure to the canonical low-conducting chalcogenide melt As 2 S 3 , we observed both the FST and the SC-M transition, occurring simultaneously within the same domain of the P, T−phase space. These findings suggest that the FST is not limited to a few exceptional liquids but is a common phenomenon, at least in systems that exhibit melt metallization within specific regions of their P, T−phase diagrams.

first-principles molecular dynamics↗

Atomic Structure, Dynamics, Changes in Chemical Bonding and Semiconductor-Metal Transition in Sb 2 Se 3 : A Remarkable Material for Quantum Networks and Energy Applications

Antimony sesquiselenide has become an outstanding functional material for photovoltaics, energy storage and transformation, memory and photonic applications. Sb 2 Se 3 is one of the most successful emerging solar light absorbers and has also been identified as a highly promising ultralow-loss phase-change material (PCM) for next-generation coherent nanophotonic processors, photonic tensor cores, quantum and neuromorphic networks. Unlike benchmark telluride PCMs, Sb 2 Se 3 features a quasi-one-dimensional (1D) crystalline structure consisting of (Sb 4 Se 6 ) ∞ ribbons, lacks the typical PCM chemical bonding, and undergoes an extended semiconductor-metal transition above the melting point. Consequently, the origin of high optical contrast between crystalline (SET) and amorphous (RESET) logic states remains elusive and presents a significant challenge. Using high-energy X-ray diffraction and Raman spectroscopy over a wide temperature range, supported by first-principles simulations and complemented by thermal, optical and electrical measurements, as well as by 121 Sb-Mossbauer spectroscopy, the quasi-1D network of orthorhombic antimony sesquiselenide was found to undergo significant evolution in amorphous and supercooled Sb 2 Se 3 , leading to lower coordination, shorter interatomic distances and a higher p-electron density on antimony, indicating changes in chemical bonding. The observed novel Sb 2 Se 3 nanocrystalline polymorph, characterized by trigonal antimony coordination and more isolated Sb-Se ribbons, could help reduce multiple trapping defect states in the bandgap, which are typical of orthorhombic Sb 2 Se 3 , thereby enhancing the power-conversion efficiency of photovoltaic devices. Semimetallic and metallic liquid Sb 2 Se 3 exhibit a gradual transformation into a denser 2D and/or 3D network with higher antimony coordination. Localized electron states in the pseudogap are becoming extended, leading to an increase in electronic conductivity σ following the relationship σ ∝ N(E F ) 2 . Liquid Sb 2 Se 3 also appears to be strongly fragile, with a nonmonotonic change in viscosity and higher atomic mobility in the metallic liquid. Furthermore, these results explain extraordinary functionalities of Sb 2 Se 3 for photonic and energy applications.

antimony↗

Atypical phase-change alloy Ga 2 Te 3 : atomic structure, incipient nanotectonic nuclei, and multilevel writing

Emerging brain-inspired computing, including artificial optical synapses, photonic tensor cores, neuromorphic networks, etc., needs phase-change materials (PCMs) of the next generation with lower energy consumption and a wider temperature range for reliable long-term operation. Gallium tellurides with higher melting and crystallization temperatures appear to be promising candidates and enable achieving the necessary requirements. Here, using high energy X-ray diffraction and Raman spectroscopy supported by first-principles simulations, we show that vitreous g-Ga 2 Te 3 films essentially have a tetrahedral local structure and sp 3 hybridization, similar to those in the stable fcc Ga 2 Te 3 polymorph and in contrast to a vast majority of typical PCMs. Nevertheless, optical pump–probe laser experiments revealed high-contrast, fast and reversible multilevel SET-RESET transitions raising a question related to the phase change mechanism. A recently observed nanotectonic compression in bulk glassy Ga–Te alloys seems to be responsible for the PCM performance. Incipient nanotectonic nuclei, reminiscent of monoclinic high-pressure HP-Te II and rhombohedral HP-Ga 2 Te 3 , are present as minorities (2–4%) in g-Ga 2 Te 3 but are suggested to grow dramatically with increasing temperature while interacting with appropriate laser pulses. This leads to co-crystallization of HP-polymorphs amplified by a high internal local pressure reaching 4–8 GPa. The metallic HP-forms provide an increasing optical and electrical contrast, favorable for reliable PCM operations, and higher energy efficiency.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Interpretable AI forecasting for numerical relativity waveforms of quasicircular, spinning, nonprecessing binary black hole mergers

We present a deep-learning artificial intelligence model (AI) that is capable of learning and forecasting the late-inspiral, merger and ringdown of numerical relativity waveforms that describe quasicircular, spinning, nonprecessing binary black hole mergers. We used the NRHybSur3dq8 surrogate model to produce train, validation and test sets of ℓ = |m| = 2 waveforms that cover the parameter space of binary black hole mergers with mass ratios q ≤ 8 and individual spins |s$^{z}_ {(1,2)}$| ≤ 0.8. These waveforms cover the time range t ∊ [-5000 M, 130 M], where t = 0M marks the merger event, defined as the maximum value of the waveform amplitude. We harnessed the ThetaGPU supercomputer at the Argonne Leadership Computing Facility to train our AI model using a training set of 1.5 million waveforms. We used 16 NVIDIA DGX A100 nodes, each consisting of 8 NVIDIA A100 Tensor Core GPUs and 2 AMD Rome CPUs, to fully train our model within 3.5 h. Our findings show that artificial intelligence can accurately forecast the dynamical evolution of numerical relativity waveforms in the time range t ∊ [-100 M, 130 M]. Sampling a test set of 190,000 waveforms, we find that the average overlap between target and predicted waveforms is ≳99% over the entire parameter space under consideration. We also combined scientific visualization and accelerated computing to identify what components of our model take in knowledge from the early and late-time waveform evolution to accurately forecast the latter part of numerical relativity waveforms. This work aims to accelerate the creation of scalable, computationally efficient and interpretable artificial intelligence models for gravitational wave astrophysics

79 ASTRONOMY AND ASTROPHYSICS↗

Understanding Mixed Precision GEMM with MPGemmFI: Insights into Fault Resilience

Emerging deep learning workloads urgently need fast general matrix multiplication (GEMM). Thus, one of the critical features of machine-learning-specific accelerators such as NVIDIA Tensor Cores, AMD Matrix Cores, and Google TPUs is the support of mixed-precision enabled GEMM. For DNN models, lower-precision FP data formats and computation offer acceptable correctness but significant performance, area, and memory footprint improvement. While promising, the mixed-precision computation on error resilience remains unexplored. To this end, we develop a fault injection framework that systematically injects fault into the mixed-precision computation results. We investigate how the faults affect the accuracy of machine learning applications. Based on the characteristics of error resilience, we offer lightweight error detection and correction solutions that significantly improve the overall model accuracy by 75% if the models experience hardware faults. The solutions can be efficiently integrated into the accelerator's pipelines.

Fang, Bo↗

Bridging the Gap Between LLMs and LNS with Dynamic Data Format and Architecture Codesign

Deep neural networks (DNNs) have achieved tremendous success in the past few years. However, their training and inference demand exceptional computational and memory resources. Quantization has been shown as an effective approach to mitigate the cost, with the mainstream data types reduced from FP32 to FP16/BF16 and recently FP8 in the latest NVIDIA H100 GPUs. With increasingly aggressive quantization, however, the conventional floating-point formats suffer from limited precision in representing numbers around zero. Recently, NVIDIA demonstrated the potential of using a Logarithmic Number System (LNS) for the next generation of tensor cores. While LNS mitigates the hurdles in representing small numbers, in this work we observed a mismatch between LNS and the emerging Large Language Models (LLM), where LLM exhibits significant outliers when directly adopting the LNS format. In this paper, we present a data-format/architecture codesign to bright this gap. On the format side, we propose a dynamic LNS format to flexibly represent outliers at a higher precision, by exploiting asymmetry in the LNS representation and identifying outliers through a per-vector basis. On the architecture side, for demonstration, we realize the dynamic LNS format in a systolic array, which can handle the irregularity of the outliers at runtime. We implement our approach on an Alveo U280 FPGA as a prototype. Experimental results show that our design can effectively handle the outliers and resolve the mismatch between LNS and LLM, contributing to an accuracy improvement of 15.4% and 16% over the floating-point and the original LNS baselines, using four state-of-the-art LLM models. Our observation and design lay a solid foundation for the large-scale adoption of the LNS format in the next-generation deep learning hardware.

Haghi, Pouya↗

Real-time High-resolution X-Ray Computed Tomography

Computed Tomography (CT) serves as a key imaging technology that relies on computationally intensive filtering and back-projection algorithms for 3D image reconstruction. While conventional high-resolution image reconstruction (> 2K3) solutions provide quick results, they typically treat reconstruction as an offline workload to be performed remotely on large-scale HPC systems. The growing demand for post-construction AI-driven analytics and the need for real-time adjustments call for high-resolution reconstruction solutions that are feasible on local computing resources, i.e. a multi-GPU server at most. In this paper, we propose a novel approach that utilizes Tensor Cores to optimize image reconstruction without sacrificing precision. We also introduce a framework designed to enable real-time execution of end-to-end distributed image reconstruction in a multi-GPU environment. Evaluations conducted on a single Nvidia A100 and H100 GPU show performance improvements of 1.91 × and 2.15 × compared to highly optimized production libraries. Furthermore, our framework, when deployed on 8-card Nvidia A100 GPU system, demonstrates the ability to reconstruct real-world datasets into 20483 volumes (32 GB) in slightly more than one minute and 40963 volumes (256 GB) in 7 minutes.

Wu, Du↗

FTTN: Feature-Targeted Testing for Numerical Properties of NVIDIA & AMD Matrix Accelerators

FTTN is a test suite to evaluate the numerical behaviors of matrix accelerators of GPUs (NVIDIA Tensor Cores and AMD Matrix Cores) in a quick and simple setting. Matrix accelerators are heavily used in today's computationally intense applications to speed up matrix multiplications. This test suite provides a comprehensive study on the numerical behaviors of these accelerators, including support for subnormals, rounding modes, extra precision bits and FMA features. Is there

Laguna Peralta, Ignacio↗

Inference-Optimized AI and High Performance Computing for Gravitational Wave Detection at Scale

We introduce an ensemble of artificial intelligence models for gravitational wave detection that we trained in the Summit supercomputer using 32 nodes, equivalent to 192 NVIDIA V100 GPUs, within 2 h. Once fully trained, we optimized these models for accelerated inference using NVIDIA TensorRT. We deployed our inference-optimized AI ensemble in the ThetaGPU supercomputer at Argonne Leadership Computer Facility to conduct distributed inference. Using the entire ThetaGPU supercomputer, consisting of 20 nodes each of which has 8 NVIDIA A100 Tensor Core GPUs and 2 AMD Rome CPUs, our NVIDIA TensorRT-optimized AI ensemble processed an entire month of advanced LIGO data (including Hanford and Livingston data streams) within 50 s. Our inference-optimized AI ensemble retains the same sensitivity of traditional AI models, namely, it identifies all known binary black hole mergers previously identified in this advanced LIGO dataset and reports no misclassifications, while also providing a 3X inference speedup compared to traditional artificial intelligence models. We used time slides to quantify the performance of our AI ensemble to process up to 5 years worth of advanced LIGO data. In this synthetically enhanced dataset, our AI ensemble reports an average of one misclassification for every month of searched advanced LIGO data. We also present the receiver operating characteristic curve of our AI ensemble using this 5 year long advanced LIGO dataset. This approach provides the required tools to conduct accelerated, AI-driven gravitational wave detection at scale.

97 MATHEMATICS AND COMPUTING↗

Design of a Low-Cost, Submersible, Digital Holographic Microscope for in Situ Microbial Imaging

The methodologies for studying marine microbiology typically consist of utilizing instrumentation within a laboratory. This typically requires extracting a sample from its place of origin prior to examination, which may be days after collection. Oftentimes, the solution is to bring the lab to the ocean, which may be costly and provide further limitations for a sterile and stable laboratory environment. Here we present a low-cost, submersible, digital holographic microscope (DHM) designed to image marine microorganisms (such as bacteria and plankton) in their natural underwater environment. Our instrument eliminates the need to transport samples and allows for instantaneous data collection of microbes in-situ. The DHM achieves sub-micron spatial resolution and is paired with artificial intelligence for the detection and tracking of specimens to reduce the overall collected data. This instrument also aims to reduce the cost of manufacturing and field expenses relative to marine microbiological research. “Off the shelf” components were selected in the design process of this instrument which allows us to achieve precise results without sacrificing data quality. The DHM itself costs under one thousand dollars and features a low-cost high-resolution camera, the Arducam MT9J001. Included in the design were five main subsystems: optical, mechanical, electrical, power, and machine learning. Our on board computer and artificial intelligence consist of a Raspberry Pi 4 (8 Gb) and Google Coral USB Tensorflow accelerator. Instrument testing has successfully proven our abilities of data acquisition for at least two hours in depths of at least forty meters below sea level. Additionally, our artificial intelligence system is currently capable of tracking up to ten areas of interest in a fraction of a second with over ninety percent confidence via the neural net driven by the tensor cores on the Google Coral. Furthermore, we have demonstrated the versatility of our instrument by mounting it on an ocean-going ROV, the BlueRov2 by BlueRobotics. Our tests on the BlueRov2 exemplified the cost-effective nature of a submersible and reusable instrument that can be implemented in moderate environments and on most vessels.

Wallace, James Kent↗

Kinematics and dynamics of disclination lines in three-dimensional nematics

An exact kinematic law for the motion of disclination lines in nematic liquid crystals as a function of the tensor order parameter Q is derived. Unlike other order parameter fields that become singular at their respective defect cores, the tensor order parameter remains regular. Following earlier experimental and theoretical work, the disclination core is defined to be the line where the uniaxial and biaxial order parameters are equal, or equivalently, where the two largest eigenvalues of Q cross. This allows an exact expression relating the velocity of the line to spatial and temporal derivatives of Q on the line, to be specified by a dynamical model for the evolution of the nematic. By introducing a linear core approximation for Q, analytical results are given for several prototypical configurations, including line interactions and motion, loop annihilation, and the response to external fields and shear flows. Behaviour that follows from topological constraints or defect geometry is highlighted. Finally, the analytic results are shown to be in agreement with three-dimensional numerical calculations based on a singular Maier–Saupe free energy that allows for anisotropic elasticity.

36 MATERIALS SCIENCE↗

Entanglement structures in quantum field theories: Negativity cores and bound entanglement in the vacuum

Here, the many-body entanglement between two finite (size-d) disjoint vacuum regions of noninteracting lattice scalar field theory in one spatial dimension, i.e., a (d A × d B ) mixed Gaussian continuous variable system, is locally transformed into a tensor-product core of (1 A × 1 B ) mixed entangled pairs. Accessible entanglement within these core pairs exhibits an exponential hierarchy and as such identifies the structure of dominant region modes from which vacuum entanglement could be extracted into a spatially separated pair of quantum detectors. Beyond the core, the remaining modes of the halo are determined to be AB separable in isolation, as well as separable from the core. However, state preparation protocols that distribute entanglement in the form of (1 A × 1 B ) mixed core pairs are found to require additional entanglement in the halo that is obscured by classical correlations. This inaccessible (bound) halo entanglement is found to mirror the accessible entanglement, but with a step behavior as the continuum is approached. It remains possible that alternate initialization protocols that do not utilize the exponential hierarchy of core-pair entanglement may require less inaccessible entanglement. Entanglement consolidation is expected to persist in higher dimensions and may aid classical and quantum simulations of asymptotically free gauge field theories, such as quantum chromodynamics.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

An Incremental Tensor Train Decomposition Algorithm

We present a new algorithm for incrementally updating the tensor train decomposition of a stream of tensor data. This new algorithm, called the tensor train incremental core expansion (TT-ICE) improves upon the current state-of-the-art algorithms for compressing in tensor train format by developing a new adaptive approach that incurs significantly slower rank growth and guarantees compression accuracy. This capability is achieved by limiting the number of new vectors appended to the TT-cores of an existing accumulation tensor after each data increment. These vectors represent directions orthogonal to the span of existing cores and are limited to those needed to represent a newly arrived tensor to a target accuracy. We provide two versions of the algorithm: TT-ICE and TT-ICE accelerated with heuristics (TT-ICE*). Here, we provide a proof of correctness for TT-ICE and empirically demonstrate the performance of the algorithms in compressing large-scale video and scientific simulation datasets. Compared to existing approaches that also use rank adaptation, TT-ICE* achieves 57× higher compression and up to 95% reduction in computational time.

97 MATHEMATICS AND COMPUTING↗

Overview of Hydraulic Fracturing Test Site 2 in the Permian Delaware Basin (HFTS-2)

Here, the Hydraulic Fracturing Test Site 2 (HFTS-2) is a large collaborative field-based R&D program in the Permian Delaware Basin, funded by the US Department of Energy through the National Energy Technology Laboratory (NETL) and the E&P industry, with support from academia. The projects' main objective is to improve the understating of the hydraulic fracturing process through utilization of advanced diagnostics and collection of through-fracture cores to provide undisputable evidence and attributes of the created hydraulic fractures. At the HFTS-2, in excess of $30 million was used to perform hydraulic fracturing research focusing on the Wolfcamp formation at a field site hosted and operated by Occidental. In addition to the research data collected by the project, Occidental provided a significant amount of background data for about a dozen existing wells in the test area as well as access to previously collected core. Additional technical and laboratory support was provided by the program members. Building on learnings and unanswered questions from HFTS-1 in the Permian Midland basin, the HFTS-2 used eight new producing wells and two existing (parent) wells to perform hydraulic fracturing research. Multiple science wells were drilled to sample and characterize the subsurface, including the collection of 540 feet of core in a vertical pilot hole and 948 feet of high-angle through-fracture core. The project installed permanent fiber optic cables in 3 wells to monitor near wellbore signals during fracturing and to collect cross-well strain measurements. Additional advanced diagnostics included a significant formation evaluation program on the vertical whole core, multi array moment tensor inversion capable microseismic survey, multi-well time-lapse geochemistry analysis, analysis of proppant distribution in producing child and slant core well, and others. We will provide an overview of the HFTS-2 project, including list of the consortium members, details of the test site, experiments performed, and technologies tested.

58 GEOSCIENCES↗

Three-dimensional atomic structure and local chemical order of medium- and high-entropy nanoalloys

Medium- and high-entropy alloys (M/HEAs) mix several principal elements with near-equiatomic composition and represent a model-shift strategy for designing previously unknown materials in metallurgy, catalysis and other fields. One of the core hypotheses of M/HEAs is lattice distortion, which has been investigated by different numerical and experimental techniques. However, determining the three-dimensional (3D) lattice distortion in M/HEAs remains a challenge. Moreover, the presumed random elemental mixing in M/HEAs has been questioned by X-ray and neutron studies, atomistic simulations, energy dispersive spectroscopy and electron diffraction, which suggest the existence of local chemical order in M/HEAs. However, direct experimental observation of the 3D local chemical order has been difficult because energy dispersive spectroscopy integrates the composition of atomic columns along the zone axes and diffuse electron reflections may originate from planar defects instead of local chemical order. Here, in this work, we determine the 3D atomic positions of M/HEA nanoparticles using atomic electron tomography and quantitatively characterize the local lattice distortion, strain tensor, twin boundaries, dislocation cores and chemical short-range order (CSRO). We find that the high-entropy alloys have larger local lattice distortion and more heterogeneous strain than the medium-entropy alloys and that strain is correlated to CSRO. We also observe CSRO-mediated twinning in the medium-entropy alloys, that is, twinning occurs in energetically unfavoured CSRO regions but not in energetically favoured CSRO ones, which represents, to our knowledge, the first experimental observation of correlating local chemical order with structural defects in any material. We expect that this work will not only expand our fundamental understanding of this important class of materials but also provide the foundation for tailoring M/HEA properties through engineering lattice distortion and local chemical order.

36 MATERIALS SCIENCE↗

Nonnegative canonical tensor decomposition with linear constraints: nnCANDELINC

Abstract There is an emerging interest for tensor factorization applications in big‐data analytics and machine learning. To speed up the factorization of extra‐large datasets, organized in multidimensional arrays (also known as tensors), easy to compute compression‐based tensor representations, such as, Tucker and tensor train formats, are used to approximate the initial large‐tensor. Further, tensor factorization is used to extract latent features that can facilitate discoveries of new mechanisms and signatures hidden in the data, where the explainability of the latent features is of principal importance. Nonnegative tensor factorization extracts latent features that are naturally sparse and parts of the data, which makes them easily interpretable. However, to take into account available domain knowledge and subject matter expertise, often additional constraints need to be imposed, which lead us to canonical decomposition with linear constraints (CANDELINC), a canonical polyadic decomposition with rank deficient factors. In CANDELINC, Tucker compression is used as a preprocessing step, which lead to a larger residual error but to more explainable latent features. Here, we propose a nonnegative CANDELINC (nnCANDELINC) accomplished via a specific nonnegative Tucker decomposition; we refer to as minimal or canonical nonnegative Tucker. We derive several results required to understand the specificity of nnCANDELINC, focusing on the difficulties of preserving the nonnegative rank of a tensor to its Tucker core and comparing the real valued to nonnegative case. Finally, we demonstrate nnCANDELINC performance on synthetic and real‐world examples.

97 MATHEMATICS AND COMPUTING↗