Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “kernel”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

Sparse Cholesky factorization for solving nonlinear PDEs via Gaussian processes

In recent years, there has been widespread adoption of machine learning-based approaches to automate the solving of partial differential equations (PDEs). Among these approaches, Gaussian processes (GPs) and kernel methods have garnered considerable interest due to their flexibility, robust theoretical guarantees, and close ties to traditional methods. They can transform the solving of general nonlinear PDEs into solving quadratic optimization problems with nonlinear, PDE-induced constraints. However, the complexity bottleneck lies in computing with dense kernel matrices obtained from pointwise evaluations of the covariance kernel, and its partial derivatives, a result of the PDE constraint and for which fast algorithms are scarce. The primary goal of this paper is to provide a near-linear complexity algorithm for working with such kernel matrices. We present a sparse Cholesky factorization algorithm for these matrices based on the near-sparsity of the Cholesky factor under a novel ordering of pointwise and derivative measurements. The near-sparsity is rigorously justified by directly connecting the factor to GP regression and exponential decay of basis functions in numerical homogenization. We then employ the Vecchia approximation of GPs, which is optimal in the Kullback-Leibler divergence, to compute the approximate factor. This enables us to compute ϵ-approximate inverse Cholesky factors of the kernel matrices with complexity O(N log d (N/ϵ)) in space and O(N log 2d (N/ϵ)) in time. We integrate sparse Cholesky factorizations into optimization algorithms to obtain fast solvers of the nonlinear PDE. We numerically illustrate our algorithm’s near-linear space/time complexity for a broad class of nonlinear PDEs such as the nonlinear elliptic, Burgers, and Monge-Ampère equations. In summary, we provide a fast, scalable, and accurate method for solving general PDEs with GPs and kernel methods.

97 MATHEMATICS AND COMPUTING↗

Numerical and Experimental Study of an Aircraft Igniter Plasma Jet Discharge

The spark discharge of an aircraft plasma jet igniter is studied using high-fidelity numerical simulations and X-ray radiography measurements. The target problem here features the thermal expansion of hot gas introduced by the electric spark within a confined igniter cavity, which eventually evolves into a pulsed jet of a high-temperature kernel. A comprehensive set of models adapted from existing strategies for internal combustion engine spark plug discharge is extended to the target problem, including the modeling of energy deposition, plasma reactions, thermodynamic properties, and heat losses. A series of validation and parameter studies are performed and presented. The kernel size is found to be sensitive to heat losses arising from radiation and hot gas remained within the discharge cavity, rather than heat conduction to the wall in the discharge cavity. Depending on the enforced shape of the post-breakdown electric arc, the spark kernel can be off-centered, tilted, and considerably asymmetric. These features have been previously not considered when studying such igniter configurations and may have a first-order impact on the ignition process. Provided a proper setup of the heat loss models and electric arc shape, the numerical results are quantitatively comparable to the experimental results in terms of the kernel size, shape, and velocity throughout different stages after the spark discharge.

Tang, Yihao↗

AGR-3/4 TRISO Fuel Compact Ceramography

The combined third and fourth irradiation in the Advanced Gas Reactor (AGR) program (AGR-3/4) contained tristructural isotropic (TRISO)-coated particle fuel and designed-to-fail (DTF) fuel particles. The DTF particles were only coated with highly anisotropic pyrocarbon (PyC) so they would purposely fail during the AGR-3/4 irradiation and provide a source of fission products for measurement. To observe the post-irradiation morphology of these DTF particles and the TRISO-coated particles, three AGR-3/4 compacts were mounted in epoxy, sectioned above their centerlines, and ground/polished. Three rounds of grinding/polishing and optical microscopy were performed so that the particles could be observed at multiple planes. Each compact contained approximately 1,872 TRISO particles and exactly 20 DTF particles. The three AGR-3/4 compacts examined covered a wide temperature range and featured both the hottest average irradiation temperature (1375°C) and the coldest/lowest burnup (872°C and 5.5% fissions per initial metal atom) of any compacts to undergo post-irradiation examination (PIE) in the AGR program to date. A total of 29 DTF particles were located and observed via microscopy across the three compacts. All the DTF particles observed via microscopy had completely failed PyC coatings. Some different DTF kernel and PyC morphologies were observed that have been attributed to the differences in irradiation temperature and/or burnup. At low and medium irradiation temperatures (roughly 850 to 1050°C), it appears irradiation-induced dimensional changes in the DTF PyC caused it to fracture and fold in on itself, and the kernel deformed to accommodate this PyC deformation. The extent of DTF kernel deformation to accommodate DTF PyC buckling tended to be greater for the medium-temperature fuel compared to the cold fuel. In the high-irradiation-temperature compact (1375°C), the DTF PyC layer appears to have completely reacted chemically with the kernel material. The only material surrounding the DTF particles in this hot compact resembles the compact matrix material. Observations of the TRISO-coated driver fuel particles and fuel compact graphitic matrix were also made. The TRISO particle kernels in the medium-temperature compact (1047°C), and many in the high-temperature compact (1375°C), showed morphologies consistent with what was seen in AGR-1 and AGR-2. In other TRISO particles from the high-temperature compact, a spatial gradient in the kernel appearance was evident. Here the cool side of the kernel (away from the center of the compact) had a much darker appearance (like that of the buffer and pyrocarbon). This is the first observation of this kind of kernel spatial gradient in AGR fuel. It is believed that the high irradiation temperature of that compact (an average of 1375°C) activates or accelerates an unidentified chemical reaction, and the temperature gradient within the fuel causes a sharp spatial gradient. The low-temperature/low-burnup compact also showed unique kernel morphologies where the TRISO-coated particles had clearly distinguishable oxide and carbide phases at the center of the kernel, an oxide rind surrounding this, and remnants of a carbide skin surrounding the oxide rind. These are features observed in the as-fabricated fuel that are no longer present in higher burnup fuel. Occasional gaps between the outer pyrolytic carbon (OPyC) and graphitic matrix material were observed in all compacts. Gaps were most often found in the small, matrix-filled spaces between adjacent particles. Finally, buffer morphologies in AGR-3/4 TRISO particles were like those observed in AGR-1 and AGR-2, and the buffer fracture frequencies in AGR-3/4 and AGR-2 TRISO particles were plotted versus fast neutron fluence and irradiation temperature. In AGR-3/4 the buffer fracture rate was highest (23%) for an irradiation temperature of 1047°C and a fast fluence of 5.18E25 n/m 2 . Increasing the irradiation temperature to 1375°C reduced the buffer fracture rate to 14%. Reducing the irradiation temperature to 872°C and the fast fluence by a factor of 3 reduced the buffer fracture rate to 7%. The temperature and fluence dependencies observed for AGR-3/4 buffer fracture were consistent with those observed in AGR-2. Lower fluences and/or higher temperatures significantly reduced the buffer fracture frequencies. Lower buffer fracture frequencies were also observed at lower temperatures as long as the fluence was also significantly reduced. This is believed to be due to lower fluence resulting in less irradiation-induced dimensional change (shrinkage) of the buffer, and a higher temperature promoting enhanced creep relaxation of stresses within the buffer, leading to less buffer fracture.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Noise and error analysis and optimization in particle-based kinetic plasma simulations

In this paper we analyze the noise in macro-particle methods used in plasma physics and fluid dynamics, leading to approaches for minimizing the total error, focusing on electrostatic models in one dimension. We begin by describing kernel density estimation for continuous values of the spatial variable x, expressing the kernel in a form in which its shape and width are represented separately. The covariance matrix of the noise in the density is computed, first for uniform true density. The bandwidth of the covariance matrix C(x,y) is related to the width of the kernel. A feature that stands out is the presence of constant negative terms in the elements of the covariance matrix both on and off-diagonal. These negative correlations are related to the fact that the total number of particles is fixed at each time step; they also lead to the property ∫C(x,y)dy = 0. We investigate the effect of these negative correlations on the electric field computed by Gauss's law, finding that the noise in the electric field is related to a process called the Ornstein-Uhlenbeck bridge, leading to a covariance matrix of the electric field with variance significantly reduced relative to that of a Brownian process. For non-constant density, p(x), still with continuous x, we analyze the total error in the density estimation and discuss it in terms of bias-variance optimization (BVO). For some characteristic length l, determined by the density and its second derivative, and kernel width h, having too few particles within h leads to too much variance; for h that is large relative to l, there is too much smoothing of the density. The optimum between these two limits is found by BVO. For kernels of the same width, it is shown that this optimum (minimum) is weakly sensitive to the kernel shape. Next, we repeat the analysis for x discretized on a grid. In this case the charge deposition rule is determined by a particle shape. An important property to be respected in the discrete system is the exact preservation of total charge on the grid; this property is necessary to ensure that the electric field is equal at both ends, consistent with periodic boundary conditions. We find that if the particle shapes satisfy a partition of unity property, the particle charge deposited on the grid is conserved exactly. Further, if the particle shape is expressed as the convolution of a kernel with another kernel that satisfies the partition of unity, then the particle shape obeys the partition of unity. This property holds for kernels of arbitrary width, including widths that are not integer multiples of the grid spacing. Furthermore, we show results relaxing the approximations used to do BVO optimization analytically, by doing numerical computations of the total error as a function of the kernel width, on a grid in x. The comparison between numerical and analytical results shows good agreement over a range of particle shapes. We discuss the practical implications of our results, including the criteria for design and implementation of computationally efficient particle shapes that take advantage of the developed theory.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

AGR-5/6/7 Fuel Fabrication Report

The U.S. Department of Energy Office of Nuclear Energy (DOE NE) and the Idaho National Laboratory (INL) Advanced Reactor Technologies (ART) Advanced Gas Reactor (AGR) Fuel Development and Qualification program (referred to as AGR Fuel program hereafter) are pursuing qualification of tristructural isotropic (TRISO) coated particle fuel for use in high temperature gas cooled reactors (HTGRs). The AGR Fuel program was established to provide a fuel qualification data set in support of the licensing and operation of an HTGR. BWX Technologies Nuclear Operations Group (BWXT-NOG) was subcontracted to fabricate the fuel for the AGR program. Several investments and innovations were realized in preparation to fabricate fuel for the AGR-5/6/7 irradiation experiments that brought fuel fabrication fully out of the laboratory and into engineering-scale operations. These included: • Increased the kernel fabrication line capacity and uniformity • Upgraded ancillary support equipment and processes for the tristructural isotropic (TRISO) coating furnace • Demonstrated efficient production of the matrix precursor powder by dry jet milling of co mingled components • Demonstrated an engineering-scale method for quick and efficient overcoating TRISO particles with the matrix precursor • Demonstrated an automated, multi cavity compacting system with a volumetric feed system • Demonstrated a combined-cycle thermal treatment furnace These changes increased production rates of some of these processes by an order of magnitude or more while eliminating the use of flammable solvents, multiple grinding and sorting operations, and the weighing out of individual die charges. Three fuel kernel lots were fabricated for production of the fuel for AGR-5/6/7. The initial lot (J52R-16-39316) was certified to fuel specifications but was not used because the kernels had a high fraction of internal fissures that caused an unacceptable fraction of the kernels to fragment when charged to the coating furnace where the TRISO coating would be deposited. Fragmented kernels increased the dispersed uranium in the particles and produced a worrisome fraction of dimpled particles with an elevated probability of in-pile failure. After some efforts to identify the cause of the fissure formation, two additional lots were produced with much lower fissure fractions, J52R-16-69317 and 69318. The latter kernel lot was a backup to the first and was not needed. Multiple kernel batches were composited to form each of the lots so as to simulate a commercial-scale operation where kernel batches would also be composited. Multiple TRISO coating runs were performed and the product characterized so that several could be composited into a TRISO lot. TRISO lot J52R-16-98005 conformed to all fuel specifications except the mean outer pyrocarbon (OPyC) layer thickness was thinner than specified. Furthermore, it was determined that the TRISO lot had a dispersed uranium fraction (DUF) that might result in the compacts not meeting the DUF specification. A review of the role of the OPyC layer and consequences of the DUF by the Technical Coordination Team and INL concluded that the fuel was acceptable for use in the AGR-5/6/7 irradiation experiment. The TRISO particles were overcoated with the matrix precursor that had been produced in a jet-milling operation. The overcoating was performed in equipment originally designed to coat pharmaceuticals. The overcoating process performed well; producing highly spherical and uniform overcoats requiring little upgrading and no recycle or rework. TRISO particles were overcoated with the matrix precursor to achieve nominal volumetric packing fractions (PFs) of TRISO particles of 25% and 40% for the irradiation experiments. The 40% PF compacts occupy the first and fifth test capsule in the test train while the inner three capsules are loaded with 25% PF compacts. The resinated graphite matrix precursor powder was a derivative of the German A3-27 matrix formulation, which differs from previous AGR irradiation campaigns that used an A3-3 formulation. Jet milling of the matrix powder precursor produced a finer mean graphite particle size than the milling operations used for the A3-3 matrix powder precursor. Changes made in the matrix formula and equipment yielded compacts with significantly higher matrix density than was attained in previous AGR irradiation campaigns. The changes in the matrix formulation and the means of milling the powders also complicated resolution of the three fuel compact defect fractions, DUF, exposed kernel fraction (EKF), and the silicon carbide defect fraction (SDF). Characterization data from BWXT-NOG had some anomalous results, so samples of the fuel compacts and overcoated TRISO particles were also analyzed by Oak Ridge National Laboratory (ORNL) to ensure that the defect fractions were accurately characterized.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

X-ray Computed Tomography of Irradiated and Unirradiated AGR-3/4 Compacts

X-ray Computed Tomography (XCT) has been utilized to image and characterize compacts from the combined third and fourth irradiation of the Advanced Gas Reactor (AGR) Program, AGR-3/4, fuel. The experiment contained tristructural isotropic (TRISO)-coated fuel particles as well as designed-to-fail (DTF) fuel particles. Two irradiated compacts, representing the lower and higher range of AGR-3/4 burnup (4.85% and 14.92% fissions per initial heavy metal atom FIMA) were examined. These represent the first known highly irradiated TRISO fuel compacts to be examined via X-ray CT. Additionally, two unirradiated compacts from the same production batch as the examined irradiation compacts were also imaged for a baseline comparison. As XCT of irradiated TRISO compacts is not a commonly implemented characterization technique, a significant portion of the report focuses on developed methodology and imaging conditions. A specialized sample shielding device was developed and fabricated specifically to limit received dose to staff during sample preparation for XCT and to minimize excess gamma radiation dose to sensitive electronic components with the utilized X-ray system. Significant penetration through the uranium oxycarbide fuel kernels by significantly hardening the X-ray beam with specialized proprietary filters acquired from Carl Zeiss NTS Ltd. The filter utilized resulted in an average X ray photon energy of ~110 keV which approaches uranium’s K-edge (~115 keV), maximizing penetration for a microfocus X-ray source. The gamma-radiation emitted from the irradiated AGR-3/4 TRISO compacts, has the same properties and mechanisms for interaction with matter as X-rays, thus the detection of gamma-radiation by the utilized X-ray detectors was initially a concern. However, although ?-rays did produce an observable signal on the X-ray detector, its contribution to the overall imaging results appeared negligible upon 3D reconstruction. The neglibile impact on the resulting 3D reconstructed volumes were likely the result of: (1) a significantly lower detection efficiency for ?-rays relative to X-rays; (2) An X-ray flux at the detector several orders of magnitude higher than that of the impinging ?-rays from the irradiated compacts. These results suggest that irradiated compacts with significantly higher radiation fields can be examined in the future if an acceptable route for sample handling and preparation can be determined. Additionally, the 3D imaging results of XCT can provide a valuable means of assessing compacts. While in many ways complimentary to traditional post irradiation examination techniques such as optical ceramography, XCT can provide additional insight into compact features traditionally difficult to discern directly from cross-sectional imaging alone. Preliminary analyses on kernel size, morphology (aspect ratio and sphericity), and kernel orientation were presented. Sphericity, a simple morphological shape descriptor, was utilized to screen for kernel extrusions within the high burnup compact. The number of kernel extrusions identified via XCT represented an approximate two-fold increase from the quantity of extruded particles observed (via optical ceramography) in adjacent compacts from the same irradiation capsule. While numerical analysis of the compact datasets was highly preliminary, initial results show promise for providing complimentary metrics to current AGR-3/4 PIE and potentially additional insight into the processes driving TRISO fuel degradation during reactor operation. Additional analyses to be performed at a later date include a more detailed examination of kernel size, kernel sphericity (and observed kernel extrusions), and sphericity. Given all particles can be observed in a single data volume possible correlation of spatial position with observed kernel features will also be made at a later date.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Adrastea: An Efficient FPGA Design Environment for Heterogeneous Scientific Computing and Machine Learning

We present Adrastea, an efficient FPGA design environment for developing scientific machine learning applications. FPGA development is challenging, from deployment, proper toolchain setup, programming methods, interfacing FPGA kernels, and more importantly, the need to explore design space choices to get the best performance and area usage from the FPGA kernel design. Adrastea provides an automated and scalable design flow to parameterize, implement, and optimize complex FPGA kernels and associated interfaces. We show how virtualization of the development environment via virtual machines is leveraged to simplify the setup of the FPGA toolchain while deploying the FPGA boards and while scaling up the automated design space exploration to leverage multiple machines concurrently. Adrastea provides an automated build and test environment of FPGA kernels. By exposing design space hyper-parameters, Adrastea can automatically search the design space in parallel to optimize the FPGA design for a given metric, usually performance or area. Adrastea simplifies the task of interfacing with the FPGA kernels with a simplified interface API. To demonstrate the capabilities of Adrastea, we implement a complex random forest machine learning kernel with 10,000 input features while achieving extremely low computing latency without loss of prediction accuracy, which is required by a scientific edge application at SNS. We also demonstrate Adrastea using an FFT kernel and show that for both applications Adrastea is able to systematically and efficiently evaluate different design options, which reduced the time and effort required to develop the kernel from months of manual work to days of automatic builds.

Young, Aaron↗

Mneme

A simple tool allowing recording the execution of a GPU (CUDA) kernel and replaying that kernel as an independent executable. The tool operates in 3 phases. During compile time the user needs to apply a provided LLVM pass to instrument the code. The pass detects all device global variables and device functions and stores this information with the respective LLVM-IR in the global device memory. The compilation generates a record-able executable. The second phase involves running the application executable with a desired input and using LD_PRELOAD to enable recording. When recording before invoking a device kernel the pre-loaded library stores device memory in persistent storage and associates the memory with the device kernel and an LLVM IR file. At the end of the recorded execution the pre-load library generates a database in the form of a JSON file containing information regarding the LLVM-IR files and the snapshots of device memory. During the third and last phase the user can replay the execution of an kernel as a separate independent executable. Besides executing it the user can modify the LLVM IR file and auto-tune parameters such as kernel launch-bounds or kernel runtime execution parameters (e.g. Kernel Block and Grid Dimensions). Is

Parasyris, Konstantinos↗

Performance Analysis of PIConGPU: Particle-in-Cell on GPUs using NVIDIA’s NSight Systems and NSight Compute

PIConGPU, Particle In Cell on GPUs, is an open source simulations framework for plasma and laser-plasma physics used to develop advanced particle accelerators for radiation therapy of cancer, high energy physics and photon science. While PIConGPU has been optimized for at least 5 years to run well on NVIDIA GPU-based clusters, there has been limited exploration by the development team of potential scalability bottlenecks using recently updated and new tools including NVIDIA’s NVProf tool and the brand-new NVIDIA NSight Suite (Systems and Compute) tools. PIConGPU is a highly optimized application that runs production jobs at scale on a system Oak Ridge Leadership Facility’s (OLCF) Summit supercomputer (using the full machine at 4600 nodes; at 98% of GPU utilization on all ~28000 NVIDIA Volta GPUs). PIConGPU has been selected as one of the the eight applications for OLCF’s coveted Center for Accelerated Application Readiness (CAAR) program aimed at the facility’s Frontier supercomputer (OLCF’s first exascale system to launch in 2021), to partner with our vendors (primary vendors: AMD and Cray/HPE) ensuring that Frontier will be able to perform large-scale science when it opens to users in 2022. To this effect, performance engineers on the PIConGPU team wanted to dive deep into the application to understand at the finest granularity, which portions of the code could be further optimized to exploit the hardware on Summit at it’s maximum potential and also to elucidate which key kernels should be tracked and optimized for the CAAR effort to port this code to Frontier. Any bottlenecks that are observed via performance profiling on Summit are likely to also impact scalability on the Frontier-dev system and the Frontier Early Access (EA) system. Additionally, the engineers wanted to take a closer look at the newest NVIDIA profiling tools which allows us to identify the most useful features on these tools and will provide an opportunity to compare it to new AMD and Cray’s performance analysis tool releases and provide feedback to our vendor partners on what features are most important and mission critical for CAAR efforts. The primary goal of this report is to focus on the evaluation of PIConGPU’s most time-intensive kernels using NVProf and NSight Suite. Three kernels, Current Deposition (also known as Compute Current), Particle Push (Move and Mark), and Shift Particles are known to be some of the most time-consuming kernels in PIConGPU. The Current xi Deposition kernel and Particle Push kernel both set up the particle attributes for running any physics simulation with PIConGPU, so it is crucial to improve the performance of these two kernels. In this report, we measure single GPU metrics for the three kernels, offer high level takeaways from the conducted analysis, and compare the profiling data from NSight Compute to that of NVProf. This analysis was performed using a grid size of 240 x 272 x 224, and 10 time steps with the Mid-November Figure of Merit (FOM) run setup. The Traveling Wave Electron Acceleration (TWEAC) science case used in this run is a representative science case for PIConGPU. This execution can also be used for baseline analysis on AMD MI50/ MI60 systems. As of the time of writing, the PIConGPU application has limited use for features of NSight Systems, so this report will mainly focus on insights garnered from NSight Compute. For this analysis, we run the “full” metric set available in NSight Compute version 2020.1.2 and use NSight Systems version 2020.3.1 to generate the application timeline.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Modeling of the effects of non-equilibrium excitation and electrode geometry on H 2 /air ignition in a nanosecond plasma discharge

In this work, we present the results of two-dimensional modeling of the effects of non-equilibrium excitation and electrode geometry on H 2 /air ignition in a nanosecond plasma discharge. A multiscale adaptive reduced chemistry solver for plasma assisted combustion (MARCS-PAC) based on PASSKEy discharge modeling package and compressible multi-component reactive flow solver ASURF+ is developed and validated. This model is applied to simulate the impact of non-equilibrium plasma excitation and electrode geometry and heat loss on the dynamics of the discharge from streamer to spark and ignition kernel development in a H 2 /air mixture with a pair of cylindrical electrodes. The results show that the plasmagenerated species (N 2 (A), N 2 (B), N 2 (a'), N 2 (C), O( 1 D), O and H) in the spark and afterglow significantly accelerate the ignition kernel development. The increase of discharge voltage at the same total discharge energy promotes the non-equilibrium active species production. It is found that the production of electronically excited species at higher reduced electric field strength is more efficient in enhancing ignition in comparison to the vibrational excitation and heating. Moreover, the 2D simulation clearly reveals that the electric field and active species distribution are highly non-uniform. The streamers are initiated at the sharp outer edges of the negative and positive electrodes by a strong electric field while the electric field is much weaker at the centerline of the electrodes. Furthermore, the simulations reveal that the ignition enhancement is sensitive to the variation of electrode shape, diameter, and gap size due to the changes of electric field distribution and location of streamer formation. A cylindrical electrode produces a larger discharge volume and ignition kernel than the parabolic and spherical electrodes, when the discharge is localized near the axis of the gap. It is found that there is a non-monotonic dependence of ignition kernel size on the electrode diameter and inter-electrode distance. The increase of electrode diameter and gap size above the optimal conditions leads to the reduction of ignition kernel volume, due to the decrease of active species concentration and gas temperature. At a larger electrode surface area and electrode diameter as well as smaller electrode gap size, the heat loss to electrode plays a greater role in reducing the ignition kernel size and slowing ignition kernel development. This work provides insights and guidance to understand the kinetic enhancement of non-equilibrium plasma and the effects of electrode geometries on ignition for the optimization ignitors in advanced engines.

42 ENGINEERING↗

Improving Runtime Performance of Tensor Computations using Rust From Python

In this work, we investigate improving the runtime performance of key computational kernels in the Python Tensor Toolbox (pyttb), a package for analyzing tensor data across a wide variety of applications. Recent runtime performance improvements have been demonstrated using Rust, a compiled language, from Python via extension modules leveraging the Python C API—e.g., web applications, data parsing, data validation, etc. Using this same approach, we study the runtime performance of key tensor kernels of increasing complexity, from simple kernels involving sums of products over data accessed through single and nested loops to more advanced tensor multiplication kernels that are key in low-rank tensor decomposition and tensor regression algorithms. In numerical experiments involving synthetically generated tensor data of various sizes and these tensor kernels, we demonstrate consistent improvements in runtime performance when using Rust from Python over 1) using Python alone, 2) using Python and the Numba just-in-time Python compiler (for loop-based kernels), and 3) using the NumPy Python package for scientific computing (for pyttb kernels).

97 MATHEMATICS AND COMPUTING↗