Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “KERNELS”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

A numerical solution for two-dimensional Fredholm integral equations of the second kind with kernels of the logarithmic potential form

Two dimensional Fredholm integral equations with logarithmic potential kernels are numerically solved. The explicit consequence of these solutions to their true solutions is demonstrated. The results are based on a previous work in which numerical solutions were obtained for Fredholm integral equations of the second kind with continuous kernels.

Gabrielsen, R. E.↗

An accurate and efficient method for evaluating the kernel of the integral equation relating pressure to normalwash in unsteady potential flow

This paper describes an accurate economical method for generating approximations to the kernel of the integral equation relating unsteady pressure to normalwash in nonplanar flow. The method is capable of generating approximations of arbitrary accuracy. It is based on approximating the algebraic part of the non elementary integrals in the kernel by exponential approximations and then integrating termwise. The exponent spacing in the approximation is a geometric sequence. The coefficients and exponent multiplier of the exponential approximation are computed by least squares so the method is completely automated. Exponential approximates generated in this manner are two orders of magnitude more accurate than the exponential approximation that is currently most often used for this purpose. Coefficients for 8, 12, 24, and 72 term approximations are tabulated in the report. Also, since the method is automated, it can be used to generate approximations to attain any desired trade-off between accuracy and computing cost.

Desmarais, R. N.↗

A method for computing the kernel of the downwash integral equation for arbitrary complex frequencies

For the design of active controls to stabilize flight vehicles, which requires the use of unsteady aerodynamics that are valid for arbitrary complex frequencies, algorithms are derived for evaluating the nonelementary part of the kernel of the integral equation that relates unsteady pressure to downwash. This part of the kernel is separated into an infinite limit integral that is evaluated using Bessel and Struve functions and into a finite limit integral that is expanded in series and integrated termwise in closed form. The developed series expansions gave reliable answers for all complex reduced frequencies and executed faster than exponential approximations for many pressure stations.

Desmarais, R. N.↗

The flare kernel in the impulsive phase

The impulsive phase of a flare is characterized by impulsive bursts of X-ray and microwave radiation, related to impulsive footpoint heating up to 50 or 60 MK, by upward gas velocities (150 to 400 km/sec) and by a gradual increase of the flare's thermal energy content. These phenomena, as well as non-thermal effects, are all related to the impulsive energy injection into the flare. The available observations are also quantitatively consistent with a model in which energy is injected into the flare by beams of energetic electrons, causing ablation of chromospheric gas, followed by convective rise of gas. Thus, a hole is burned into the chromosphere; at the end of impulsive phase of an average flare the lower part of that hole is situated about 1800 km above the photosphere. H alpha and other optical and UV line emission is radiated by a thin layer (approx. 20 km) at the bottom of the flare kernel. The upward rising and outward streaming gas cools down by conduction in about 45 s. The non-thermal effects in the initial phase are due to curtailing of the energy distribution function by escape of energetic electrons. The single flux tube model of a flare does not fit with these observations; instead we propose the spaghetti-bundle model. Microwave and gamma-ray observations suggest the occurrence of dense flare knots of approx. 800 km diameter, and of high temperature. Future observations should concentrate on locating the microwave/gamma-ray sources, and on determining the kernel's fine structure and the related multi-loop structure of the flaring area.

Dejager, C.↗

On the solution of integral equations with a generalized cauchy kernel

In this paper a certain class of singular integral equations that may arise from the mixed boundary value problems in nonhomogeneous materials is considered. The distinguishing feature of these equations is that in addition to the Cauchy singularity, the kernels contain terms that are singular only at the end points. In the form of the singular integral equations adopted, the density function is a potential or a displacement and consequently the kernel has strong singularities of the form (t-x) sup-2, x sup n-2 (t+x) sup n, (n or = 2, 0x,tb). The complex function theory is used to determine the fundamental function of the problem for the general case and a simple numerical technique is described to solve the integral equation. Two examples from the theory of elasticity are then considered to show the application of the technique.

Kaya, A. C.↗

Factorization and the synthesis of optimal feedback kernels for differential-delay systems

A combination of ideas from the theories of operator Riccati equations and Volterra factorizations leads to the derivation of a novel, relatively simple set of hyperbolic equations which characterize the optimal feedback kernel for the finite-time regulator problem for autonomous differential-delay systems. Analysis of these equations elucidates the underlying structure of the feedback kernel and leads to the development of fast and accurate numerical methods for its computation. Unlike traditional formulations based on the operator Riccati equation, the gain is characterized by means of classical solutions of the derived set of equations. This leads to the development of approximation schemes which are analogous to what has been accomplished for systems of ordinary differential equations with given initial conditions.

Milman, Mark M.↗

Density Estimation with Mercer Kernels

We present a new method for density estimation based on Mercer kernels. The density estimate can be understood as the density induced on a data manifold by a mixture of Gaussians fit in a feature space. As is usual, the feature space and data manifold are defined with any suitable positive-definite kernel function. We modify the standard EM algorithm for mixtures of Gaussians to infer the parameters of the density. One benefit of the approach is it's conceptual simplicity, and uniform applicability over many different types of data. Preliminary results are presented for a number of simple problems.

Macready, William G.↗

Runtime Thread-Block Optimization for Custom Multistream CUDA Kernels for the Glenn Research Center Communication Analysis Suite

In preparation of the return of humans to the Moon with the coming Artemis missions, NASA scientists must evaluate proposed landing site locations for terrain and communications viability. The Glenn Research Center Communication Analysis Suite (GCAS) combines sophisticated communication network models with accurate lunar terrain to access sites across the Moon’s south pole. Given the importance of proper site selection to crew safety and mission success, many locations need to be analyzed resulting in a large computational load needing to be performed. To meet the growing project demands, development has begun to improve the runtime efficiency of GCAS with GPU parallelization by way of multi-stream CUDA kernels. One of the most prominent factors in kernel optimization is the proper selection of thread-block dimensions in order to maximize the concurrent operation on the device. Typically, thread-block dimensions are optimized by hand requiring many stages of benchmarking and iteration. Additionally, given the main conditions to optimization are the physical GPU architecture and problem size, these optimal dimensions are non-portable and fragile in their scope. As such, a novel optimization routine was developed to generate the optimal thread-block dimensions during runtime with considerations to hardware specifications and problem size resolving the issues of portability and enabling the function of more dynamic routines.

Aden Bergstresser↗

Systems and methods for customizing kernel machines with deep neural networks

A method including receiving an input data set. The input data set can include one of a feature domain set or a kernel matrix. The method also can include constructing dense embeddings using: (i) Nyström approximations on the input data set when the input data set comprises the kernel matrix, and (ii) clustered Nyström approximations on the input data set when the input data set comprises the feature domain set. The method additionally can include performing representation learning on each of the dense embeddings using a multi-layer fully-connected network for each of the dense embeddings to generate latent representations corresponding to each of the dense embeddings. The method further can include applying a fusion layer to the latent representations corresponding to the dense embeddings to generate a combined representation. The method additionally can include performing classification on the combined representation. Other embodiments of related systems and methods are also disclosed.

Song, Huan↗

Dynamic kernel memory space allocation

A processing unit includes one or more processor cores and a set of registers to store configuration information for the processing unit. The processing unit also includes a coprocessor configured to receive a request to modify a memory allocation for a kernel concurrently with the kernel executing on the at least one processor core. The coprocessor is configured to modify the memory allocation by modifying the configuration information stored in the set of registers. In some cases, initial configuration information is provided to the set of registers by a different processing unit. The initial configuration information is stored in the set of registers prior to the coprocessor modifying the configuration information.

Gutierrez, Anthony↗

Application of performance portability solutions for GPUs and many-core CPUs to track reconstruction kernels

Next generation High-Energy Physics (HEP) experiments are presented with significant computational challenges, both in terms of data volume and processing power. Using compute accelerators, such as GPUs, is one of the promising ways to provide the necessary computational power to meet the challenge. The current programming models for compute accelerators often involve using architecture-specific programming languages promoted by the hardware vendors and hence limit the set of platforms that the code can run on. Developing software with platform restrictions is especially unfeasible for HEP communities as it takes significant effort to convert typical HEP algorithms into ones that are efficient for compute accelerators. Multiple performance portability solutions have recently emerged and provide an alternative path for using compute accelerators, which allow the code to be executed on hardware from different vendors. We apply several portability solutions, such as Kokkos, SYCL, C++17 std::execution::par and Alpaka, on two mini-apps extracted from the mkFit project: p2z and p2r. These apps include basic kernels for a Kalman filter track fit, such as propagation and update of track parameters, for detectors at a fixed z or fixed r position, respectively. The two mini-apps explore different memory layout formats. We report on the development experience with different portability solutions, as well as their performance on GPUs and many-core CPUs, measured as the throughput of the kernels from different GPU and CPU vendors such as NVIDIA, AMD and Intel.

Kwok, Ka Martin↗

First constraints on the nonperturbative gluon Collins-Soper kernel

The gluon Collins-Soper kernel, which encodes the rapidity evolution of transverse-momentum-dependent gluon distributions, is constrained for the first time in the nonperturbative regime, for transverse momentum scales $q_{T} \in [ 300\text{ MeV}, 1.3\text{ GeV}]$. The constraints are determined in lattice QCD at a close-to-physical pion mass $M_π= 172(3)\text{ MeV}$, a single lattice spacing $a=0.15\text{ fm}$, and next-to-next-to-leading logarithmic matching in Large-Momentum Effective Theory. These results represent the first step toward a controlled determination of the gluon Collins-Soper kernel in QCD, with eventual phenomenological import and relevance to present and future experiments sensitive to the gluon structure of hadronic matter.

Avkhadiev, Artur [Argonne; MIT, Cambridge, CTP]↗

Kernel fusion in atomistic spin dynamics simulations on Nvidia GPUs using tensor core

In atomistic spin dynamics simulations, the time cost of constructing the space- and time-displaced pair correlation function in real space increases quadratically as the number of spins N, leading to significant computational effort. The GEMM subroutine can be adopted to accelerate the calculation of the dynamical spin-spin correlation function, but the computational cost of simulating large spin systems (>40000 spins) on CPUs remains expensive. In this work, we perform the simulation on the graphics processing unit (GPU), a hardware solution widely used as an accelerator for scientific computing and deep learning. Here we show that GPUs can accelerate the simulation up to 25-fold compared to multi-core CPUs when using the GEMM subroutine on both. To hide memory latency, we fuse the element-wise operation into the GEMM kernel using CUTLASS that can improve the performance by 26% ~ 33% compared to implementation based on cuBLAS. Furthermore, we perform the on-the-fly calculation in the epilogue of the GEMM subroutine to avoid saving intermediate results on global memory, which makes the large-scale atomistic spin dynamics simulation feasible and affordable.

97 MATHEMATICS AND COMPUTING↗

Exploring Causal Physical Mechanisms via Non-Gaussian Linear Models and Deep Kernel Learning: Applications for Ferroelectric Domain Structures

Rapid emergence of multimodal imaging in scanning probe, electron, and optical microscopies has brought forth the challenge of understanding the information contained in these complex data sets, targeting the intrinsic correlations between different channels, and further exploring the underpinning causal physical mechanisms. Here, we develop such an analysis framework for Piezoresponse Force Microscopy. We argue that under certain conditions, we can bootstrap experimental observations with the prior knowledge of materials structure to get information on certain nonobserved properties, and demonstrate linear causal analysis for PFM observables. We further demonstrate that the strength of individual causal links between complex descriptors can be ascertained using the deep kernel learning (DKL) model. In this DKL analysis, we use the prior information on domain structure within the image to predict the physical properties. This analysis demonstrates the correlative relationships between morphology, piezoresponse, elastic property, etc., at nanoscale. The prediction of morphology and other physical parameters illustrates a mutual interaction between surface condition and physical properties in ferroelectric materials. Overall, this analysis is universal and can be extended to explore the correlative relationships of other multichannel data sets, and allow for high-fidelity reconstruction of underpinning functionalities and physical mechanisms.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Deep kernel methods learn better: from cards to process optimization

Abstract The ability of deep learning methods to perform classification and regression tasks relies heavily on their capacity to uncover manifolds in high-dimensional data spaces and project them into low-dimensional representation spaces. In this study, we investigate the structure and character of the manifolds generated by classical variational autoencoder (VAE) approaches and deep kernel learning (DKL). In the former case, the structure of the latent space is determined by the properties of the input data alone, while in the latter, the latent manifold forms as a result of an active learning process that balances the data distribution and target functionalities. We show that DKL with active learning can produce a more compact and smooth latent space which is more conducive to optimization compared to previously reported methods, such as the VAE. We demonstrate this behavior using a simple cards dataset and extend it to the optimization of domain-generated trajectories in physical systems. Our findings suggest that latent manifolds constructed through active learning have a more beneficial structure for optimization problems, especially in feature-rich target-poor scenarios that are common in domain sciences, such as materials synthesis, energy storage, and molecular discovery. The Jupyter Notebooks that encapsulate the complete analysis accompany the article.

97 MATHEMATICS AND COMPUTING↗

Next-Cycle Optimal Dilute Combustion Control via Online Learning of Cycle-to-Cycle Variability Using Kernel Density Estimators

Dilute combustion using exhaust gas recirculation (EGR) presents a cost-effective method for increasing the efficiency of spark-ignition (SI) engines. However, the maximum amount of EGR that can be used at a given condition is limited by a rapid increment of cycle-to-cycle variability (CCV). This study describes a methodology to design a model-based stochastic optimal controller to adjust the cycle-to-cycle fuel injection quantity in order to reduce CCV and further extend the dilute limit. Given the complexity and chaotic nature of combustion events, the controller was enhanced with online learning in order to identify the statistical properties of combustion efficiency, which are needed to generate predictions for next-cycle events. This study showed that a kernel density estimator (KDE) can be used to learn the combustion properties in real time and can be incorporated into the feedback policy in order to calculate the optimal control command. Experimental results suggested that the dilute limit can be extended from 18.5% to 21% EGR fraction at an operating condition relevant for highway cruising. Additionally, the proposed controller can achieve a large CCV reduction with less fuel enrichment compared to previous methods, overall contributing to an increase in 0.2% indicated fuel conversion efficiency.

33 ADVANCED PROPULSION SYSTEMS↗

Cyber Spoofing Detection for Grid Distributed Synchrophasor Using Dynamic dual-Kernel SVM

Cyber spoofing with distributed synchrophasor adversely affects the decision-making and situational awareness of the power grid. To detect the spoofing trail, this letter proposes a composite signature-based cyber spoofing detection methodology. The intrinsic principal modes are first extracted from the distributed synchrophasor data. Then, multiple signatures of different intrinsic model components are derived to quantify the spoofing. Thereafter, the dynamic dual-kernel support vector machine is proposed to identify cyber spoofing using multiple signatures. Multiple experimental results using six spoofing methods have verified the validity of the methodology.

24 POWER TRANSMISSION AND DISTRIBUTION↗