Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “kernel”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Revealing quasi-excitations in the low-density homogeneous electron gas with model exchange–correlation kernels

Time-dependent density functional theory within the linear response regime provides a solid mathematical framework to capture excitations. The accuracy of the theory, however, largely depends on the approximations for the exchange–correlation (xc) kernels. Away from the long-wavelength (or q = 0 short wave-vector) and zero-frequency (ω = 0) limit, the correlation contribution to the kernel becomes more relevant and dominant over exchange. The dielectric function, in principle, can encompass xc effects relevant to describe low-density physics. Furthermore, besides collective plasmon excitations, the dielectric function can reveal collective electron–hole excitations, often dubbed “ghost excitons.” Besides collective excitons, the physics of the low-density regime is rich, as exemplified by a static charge-density wave that was recently found for r s > 69, and was shown to be associated with softening of the plasmon mode. These excitations are seen to be present in much higher density 2D homogeneous electron gases of r s ≳ 4. Here, in this work, we perform a thorough analysis with xc model kernels for excitations of various nature. The uniform electron gas, as a useful model of real metallic systems, is used as a platform for our analysis. We highlight the relevance of exact constraints as we display and explain screening and excitations in the low-density region.

36 MATERIALS SCIENCE↗

A fractional calculus framework for open quantum dynamics: From Liouville to Lindblad to memory kernels

Open quantum systems exhibit dynamics ranging from unitary evolution to irreversible dissipation. While the Gorini–Kossakowski–Sudarshan–Lindblad equation uniquely characterizes Markovian completely positive and trace-preserving (CPTP) evolution, many physical platforms display non-Markovian features such as algebraic relaxation and coherence backflow. Fractional calculus provides a natural way to model such long-memory behavior through power-law temporal kernels introduced by fractional time derivatives. Here, we develop a unified framework that embeds fractional master equations within the broader hierarchy of open-system formalisms. The fractional equation forms a structured subclass of memory-kernel models, reduces to the Lindblad form at unit order, and, through Bochner–Phillips subordination, admits a CPTP representation as an average over Lindblad semigroups. Its resolvent structure further connects fractional dynamics to established non-Markovian approaches, including Nakajima–Zwanzig kernels and hierarchical equations of motion, providing a compact surrogate for long-memory effects. This formulation positions fractional calculus as a rigorous and practical language for modeling non-Markovian quantum dynamics in chemical physics and physical chemistry, providing a CPTP-preserving, computationally efficient surrogate for structured condensed-phase environments where long-time memory and dissipation play a central role.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Transverse-momentum-dependent pion structures from lattice QCD: Collins-Soper kernel, soft factor, TMDWF, and TMDPDF

We present the first lattice quantum chromodynamics (QCD) calculation of the pion valence-quark transverse-momentum-dependent parton distribution function (TMDPDF) within the framework of large-momentum effective theory (LaMET). Using correlators fixed in the Coulomb gauge (CG), we computed the quasi-TMD beam function for a pion with a mass of 300 MeV, a fine lattice spacing of 𝑎 =0.06 fm, and multiple large momenta up to 3 GeV. The intrinsic soft functions in the CG approach are extracted from form factors with large momentum transfer, and as a byproduct, we also obtain the corresponding Collins-Soper (CS) kernel. Our determinations of both the soft function and the CS kernel agree with perturbation theory at small transverse separations (𝑏 ⊥ ) between the quarks. At larger 𝑏 ⊥ , the CS kernel remains consistent with recent results obtained using both CG and gauge-invariant TMD correlators in the literature. By combining next-to-leading logarithmic factorization of the quasi-TMD beam function and the soft function, we obtain an 𝑥-dependent pion valence-quark TMDPDF for transverse separations 𝑏 ⊥ ≳1 fm. Interestingly, we find that the 𝑏 ⊥ dependence of the phenomenological parametrizations of TMDPDF for moderate values of 𝑥 are in reasonable agreement with our QCD determinations. In addition, we present results for the transverse-momentum-dependent wave function for a heavier pion with 670 MeV mass.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Representation Learning via Quantum Neural Tangent Kernels

Variational quantum circuits are used in quantum machine learning and variational quantum simulation tasks. Designing good variational circuits or predicting how well they perform for given learning or optimization tasks is still unclear. Here we discuss these problems, analyzing variational quantum circuits using the theory of neural tangent kernels. We define quantum neural tangent kernels, and derive dynamical equations for their associated loss function in optimization and learning tasks. We analytically solve the dynamics in the frozen limit, or lazy training regime, where variational angles change slowly and a linear perturbation is good enough. We extend the analysis to a dynamical setting, including quadratic corrections in the variational angles. We then consider a hybrid quantum classical architecture and define a large-width limit for hybrid kernels, showing that a hybrid quantum classical neural network can be approximately Gaussian. The results presented here show limits for which analytical understandings of the training dynamics for variational quantum circuits, used for quantum machine learning and optimization problems, are possible. These analytical results are supported by numerical simulations of quantum machine-learning experiments.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Automatic Generation of High-Performance Convolution Kernels on ARM CPUs for Deep Learning

In this work, we present FastConv, a template-based code auto-generation open source library that can automatically generate high-performance deep learning convolution kernels of arbitrary matrices/tensors shapes. FastConv is based on the Winograd algorithm, which is reportedly the highest performing algorithm for the time-consuming convolution layers of convolutional neural networks. ARM CPUs cover a wide range designs and specifications, from embedded devices to HPC-grade CPUs. The leads to the dilemma of how to consistently optimize Winograd-based convolution solvers for convolution layers of different shapes. FastConv addresses this problem by using templates to auto-generate multiple shapes of tuned kernels variants suitable for skinny tall matrices. As a performance portable library, FastConv transparently searches for the best combination of kernel shapes, cache tiles, scheduling of loop orders, packing strategies, access patterns, and online/offline computations. Auto-tuning is used to search the parameter configuration space for the best performance for a given target architecture and problem size. The experiments with layer-wise evaluation on the VGG--16 model confirms a 1.25x performance gains is got by tuning the Winograd library. Integrated comparison results shows 1.02x to 1.40x, 1.14x to 2.17x, and 1.22x and 2.48x speedup is achieved over NNPACK, Arm NN, and FeatherCNN on the Kunpeng 920 beside few cases. Furthermore, problem size performance portability experiments with various convolution shapes shows that FastConv achieves 1.2x to 1.7x speedup and 2x to 22x speedup over NNPACK and ARM NN inference engine using Winograd on Kunpeng 920 . CPU performance portability evaluation on the VGG--16 show an average speedup over NNPACK of 1.42x, 1.21x, 1.26x, 1.37x, 2.26x, and 11.02x is observed on Kunpeng 920, Snapdragon 835, 855, 888, Apple M1, and AWS Graviton2, respectively.

97 MATHEMATICS AND COMPUTING↗

Integral Kernel Methods for Nonlinear Parabolic-Elliptic Systems

Nonlinear parabolic-elliptic systems arise in many physical, biological, and chemical phenomena such as chemotaxis, ion transport, self-gravitating particles, and Brownian vortices. Existing methods struggle with the strong coupling and high nonlinearity and nonlocality of some of these systems, especially the ill-conditioned, convection-dominated problems. To overcome numerical difficulties, current approaches rely on initial guesses, preconditioning, or iterative techniques with no convergence guarantees. They might suffer from poor scalability, large memory usage, and difficulty to parallelize. Inspired by the connection of parabolic-elliptic systems to stochastic processes, we introduce a novel meshless, monolithic, and fully explicit method that naturally encapsulates the elliptic and parabolic operators into a single step which updates each node deterministically with global information. By being fully quadrature-based, it avoids solving systems of discretized equations and does not utilize initial guesses or preconditioning, while requiring little memory and being easy to parallelize. We first derive the method in an integral kernel formulation with quadratic complexity in the number of integration nodes and then leverage kernel-independent fast multipole methods (FMM) to present a scalable algorithm with linear complexity. We provide numerical examples for the Poisson-Nernst-Planck equations in one, two, and three dimensions, together with the derivation of the integral kernel for each case. Furthermore, the examples demonstrate the fast convergence and scalability of the FMM-accelerated algorithm, as well as its suitability for convection-dominated problems, making it competitive against traditional PDE solvers.

PDE systems↗

SW4 Curvilinear Kernels

Five computationally expensive stencil evaluation routines from SW4(https://github.com/geodynamics/sw4 GPL license) have been extracted and packaged with a driver to create a mini-app for evaluating compiler and GPU performance. The kernels can executed on AMD and Nvidia GPUs with and without RAJA. The kernel driver generates synthetic inputs, checks for correctness and measures kernel run times.

Pankajakshan, Ramesh↗

ExaSGD: 2022 Kernel Thrust Activities

The Kernel Thrust milestone ADSE22-407 covers the development of device-capable optimization algorithms and solvers technologies required by the ExaSGD project’s software stack in order to solve security-constrained alternating current optimal power flow (SC-ACOPF) problems on emerging exascale architectures. To this extent, in FY22 the main objective of the Kernel Thrust was (i) provide sparse optimization solver that runs efficiently on hardware accelerator devices (i.e., NVIDIA and AMD GPUs) to perform intra-node computations, (ii) strengthen the reliability and increase the performance of the mixed-dense sparse (MDS) solver of HiOp for deployment on the FY22 target architectures, Summit and Crusher, and (iii) increase performance by improving the mathematical algorithm and refining the parallel MPI-based implementation of the coarse-grain parallel solver HiOp-PriDec for capabilities deployment on the FY22 target architectures, Summit and Crusher. This document presents the developments and contributions done by the Kernels Thrust Team in FY22 toward completion of the above-mentioned objectives. These contributions progressed along four main development (sub)thrusts: (1) Design and implementation of a sparse optimization solver for use on hardware accelerators; (2) Improvement of the mathematical algorithm and of the parallel implementation of HiOp-PriDec to ensure readiness and efficient coarse-grain parallelism for FY23 target exascale machine; and (3) Support Software and Application Development Thrusts of the exaSGD project in their deployment of the project’s software stack on AMD- and NVIDIA-based architectures. The development of the sparse optimization solver (thrust 1 above) was new in FY22 and resulted in a new sparse solver in HiOp (available as of version 0.6). The second development thrust was a continuation of the efforts from FY21 and improved the mathematical algorithm and the communication strategy of the HiOp-PriDec solver. The last developement thrust is a large collaborative effort. Namely, the project’s teams from multiple labs (LLNL, PNNL, ORNL, and NREL) performed large-scale demonstration of the ExaSGD software stack, namely the optimization solvers of HiOp interfaced with the modeling front-end ExaGO and the stochastic sampler PowerScenarios. These demonstration efforts solved large-scale instances of the SC-ACOPF challenge problem of medium network sizes (10, 000-bus system) and large number of contingencies on Summit (NVIDIA accelerators) and Crusher (AMD accelerators) systems at ORNL.

97 MATHEMATICS AND COMPUTING↗

The Collins-Soper Kernel from Lattice QCD

I will present the first complete determination of the quark Collins-Soper kernel, which relates TMDs at different rapidity scales, using lattice QCD and including systematic control of quark mass, operator mixing, and discretization effects. Next-to-next-to-leading logarithmic matching is used to match lattice-calculable distributions to the corresponding TMDs. The continuum-extrapolated lattice QCD results are consistent with several recent phenomenological parametrizations of the Collins-Soper kernel and are precise enough to disfavor other parametrizations. I will also discuss a first exploration of the gluon Collins-Soper kernel.

Wagman, Michael [Fermilab]↗

Cholesky-based experimental design for Gaussian process and kernel-based emulation and calibration.

Gaussian processes and other kernel-based methods are used extensively to construct approximations of multivariate data sets. The accuracy of these approximations is dependent on the data used. This paper presents a computationally efficient algorithm to greedily select training samples that minimize the weighted L p error of kernel-based approximations for a given number of data. The method successively generates nested samples, with the goal of minimizing the error in high probability regions of densities specified by users. The algorithm presented is extremely simple and can be implemented using existing pivoted Cholesky factorization methods. Training samples are generated in batches which allows training data to be evaluated (labeled) in parallel. For smooth kernels, the algorithm performs comparably with the greedy integrated variance design but has significantly lower complexity. Numerical experiments demonstrate the efficacy of the approach for bounded, unbounded, multi-modal and non-tensor product densities. We also show how to use the proposed algorithm to efficiently generate surrogates for inferring unknown model parameters from data using Bayesian inference.

97 MATHEMATICS AND COMPUTING↗

Improvements to the kernel function method of steady, subsonic lifting surface theory

The application of a kernel function lifting surface method to three dimensional, thin wing theory is discussed. A technique for determining the influence functions is presented. The technique is shown to require fewer quadrature points, while still calculating the influence functions accurately enough to guarantee convergence with an increasing number of spanwise quadrature points. The method also treats control points on the wing leading and trailing edges. The report introduces and employs an aspect of the kernel function method which apparently has never been used before and which significantly enhances the efficiency of the kernel function approach.

Medan, R. T.↗

Reformulation of Possio's kernel with application to unsteady wind tunnel interference

An efficient method for computing the Possio kernel has remained elusive up to the present time. In this paper the Possio is reformulated so that it can be computed accurately using existing high precision numerical quadrature techniques. Convergence to the correct values is demonstrated and optimization of the integration procedures is discussed. Since more general kernels such as those associated with unsteady flows in ventilated wind tunnels are analytic perturbations of the Possio free air kernel, a more accurate evaluation of their collocation matrices results with an exponential improvement in convergence. An application to predicting frequency response of an airfoil-trailing edge control system in a wind tunnel compared with that in free air is given showing strong interference effects.

Fromme, J. A.↗

A solution for two-dimensional Fredholm integral equations of the second kind with periodic, semiperiodic, or nonperiodic kernels

A numerical scheme for solving two dimensional Fredholm integral equations of the second kind is developed. The proof of the convergence of the numerical scheme is shown for three cases: the case of periodic kernels, the case of semiperiodic kernels, and the case of nonperiodic kernels. Applications to the incompressible, stationary Navier-Stokes problem are of primary interest.

Gabrielsen, R. E.↗

Generalization of the subsonic kernel function in the s-plane, with applications to flutter analysis

A generalized subsonic unsteady aerodynamic kernel function, valid for both growing and decaying oscillatory motions, is developed and applied in a modified flutter analysis computer program to solve the boundaries of constant damping ratio as well as the flutter boundary. Rates of change of damping ratios with respect to dynamic pressure near flutter are substantially lower from the generalized-kernel-function calculations than from the conventional velocity-damping (V-g) calculation. A rational function approximation for aerodynamic forces used in control theory for s-plane analysis gave rather good agreement with kernel-function results, except for strongly damped motion at combinations of high (subsonic) Mach number and reduced frequency.

Cunningham, H. J.↗

On the Kernel function of the integral equation relating lift and downwash distributions of oscillating wings in supersonic flow

This report treats the Kernel function of the integral equation that relates a known or prescribed downwash distribution to an unknown lift distribution for harmonically oscillating wings in supersonic flow. The treatment is essentially an extension to supersonic flow of the treatment given in NACA report 1234 for subsonic flow. For the supersonic case the Kernel function is derived by use of a suitable form of acoustic doublet potential which employs a cutoff or Heaviside unit function. The Kernel functions are reduced to forms that can be accurately evaluated by considering the functions in two parts: a part in which the singularities are isolated and analytically expressed, and a nonsingular part which can be tabulated.

Watkins, Charles E↗

A method of smoothed particle hydrodynamics using spheroidal kernels

We present a new method of three-dimensional smoothed particle hydrodynamics (SPH) designed to model systems dominated by deformation along a preferential axis. These systems cause severe problems for SPH codes using spherical kernels, which are best suited for modeling systems which retain rough spherical symmetry. Our method allows the smoothing length in the direction of the deformation to evolve independently of the smoothing length in the perpendicular plane, resulting in a kernel with a spheroidal shape. As a result the spatial resolution in the direction of deformation is significantly improved. As a test case we present the one-dimensional homologous collapse of a zero-temperature, uniform-density cloud, which serves to demonstrate the advantages of spheroidal kernels. We also present new results on the problem of the tidal disruption of a star by a massive black hole.

Fulbright, Michael S.↗

Error and Complexity Analysis for a Collocation-Grid-Projection Plus Precorrected-FFT Algorithm for Solving Potential Integral Equations with LaPlace or Helmholtz Kernels

In this paper we derive error bounds for a collocation-grid-projection scheme tuned for use in multilevel methods for solving boundary-element discretizations of potential integral equations. The grid-projection scheme is then combined with a precorrected FFT style multilevel method for solving potential integral equations with 1/r and e(sup ikr)/r kernels. A complexity analysis of this combined method is given to show that for homogeneous problems, the method is order n natural log n nearly independent of the kernel. In addition, it is shown analytically and experimentally that for an inhomogeneity generated by a very finely discretized surface, the combined method slows to order n(sup 4/3). Finally, examples are given to show that the collocation-based grid-projection plus precorrected-FFT scheme is competitive with fast-multipole algorithms when considering realistic problems and 1/r kernels, but can be used over a range of spatial frequencies with only a small performance penalty.

Phillips, J. R.↗

KNBD: A Remote Kernel Block Server for Linux

I am developing a prototype of a Linux remote disk block server whose purpose is to serve as a lower level component of a parallel file system. Parallel file systems are an important component of high performance supercomputers and clusters. Although supercomputer vendors such as SGI and IBM have their own custom solutions, there has been a void and hence a demand for such a system on Beowulf-type PC Clusters. Recently, the Parallel Virtual File System (PVFS) project at Clemson University has begun to address this need (1). Although their system provides much of the functionality of (and indeed was inspired by) the equivalent file systems in the commercial supercomputer market, their system is all in user-space. Migrating their 10 services to the kernel could provide a performance boost, by obviating the need for expensive system calls. Thanks to Pavel Machek, the Linux kernel has provided the network block device (2) with kernels 2.1.101 and later. You can configure this block device to redirect reads and writes to a remote machine's disk. This can be used as a building block for constructing a striped file system across several nodes.

Becker, Jeff↗