Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “kernel method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Data-driven learning of nonlocal models: from high-fidelity simulations to constitutive laws

We show that machine learning can improve the accuracy of simulations of stress waves in one-dimensional composite materials. We propose a data-driven technique to learn nonlocal constitutive laws for stress wave propagation models. The method is an optimization-based technique in which the nonlocal kernel function is approximated via Bernstein polynomials. The kernel, including both its functional form and parameters, is derived so that when used in a nonlocal solver, it generates solutions that closely match high-fidelity data. The optimal kernel therefore acts as a homogenized nonlocal continuum model that accurately reproduces wave motion in a smaller-scale, more detailed model that can include multiple materials. We apply this technique to wave propagation within a heterogeneous bar with a periodic microstructure. Several one-dimensional numerical tests illustrate the accuracy of our algorithm. The optimal kernel is demonstrated to reproduce high-fidelity data for a composite material in applications that are substantially different from the problems used as training data.

97 MATHEMATICS AND COMPUTING↗

Advanced panel-type influence coefficient methods applied to unsteady three dimensional potential flows

A panel method for solving unsteady, subsonic wind-body-tail flow problems is formulated and partially verified. The method is applicable to general aircraft configurations consisting of arbitrary arrangements of wings, bodies, tails, and nacelles. The wake may be located arbitrarily and the unsteady, transverse component of vorticity in the wake may be assigned any covection velocity. The wake in the unsteady flow problem, therefore, can be given the location and convection velocity of the wake produced by a steady flow which is the mean flow of the unsteady flow problem. The panel method has been used as a basis for expanding the unsteady kernel function in a power series to obtain panel influence coefficients which can be integrated in closed form.

Dusto, A. R.↗

Factorization and the synthesis of optimal feedback kernels for differential-delay systems

A combination of ideas from the theories of operator Riccati equations and Volterra factorizations leads to the derivation of a novel, relatively simple set of hyperbolic equations which characterize the optimal feedback kernel for the finite-time regulator problem for autonomous differential-delay systems. Analysis of these equations elucidates the underlying structure of the feedback kernel and leads to the development of fast and accurate numerical methods for its computation. Unlike traditional formulations based on the operator Riccati equation, the gain is characterized by means of classical solutions of the derived set of equations. This leads to the development of approximation schemes which are analogous to what has been accomplished for systems of ordinary differential equations with given initial conditions.

Milman, Mark M.↗

Consistent Pl Analysis of Aqueous Uranium-235 Critical Assemblies

The lethargy-dependent equations of the consistent Pl approximation to the Boltzmann transport equation for slowing down neutrons have been used as the basis of an IBM 704 computer program. Some of the effects included are (1) linearly anisotropic center of mass elastic scattering, (2) heavy element inelastic scattering based on the evaporation model of the nucleus, and (3) optional variation of the buckling with lethargy. The microscopic cross-section data developed for this program covered 473 lethargy points from lethargy u = 0 (10 Mev) to u = 19.8 (0.025 ev). The value of the fission neutron age in water calculated here is 26.5 square centimeters; this value is to be compared with the recent experimental value given as 27.86 square centimeters. The Fourier transform of the slowing-down kernel for water to indium resonance energy calculated here compared well with the Fourier transform of the kernel for water as measured by Hill, Roberts, and Fitch. This method of calculation has been applied to uranyl fluoride - water solution critical assemblies. Theoretical results established for both unreflected and fully reflected critical assemblies have been compared with available experimental data. The theoretical buckling curve derived as a function of the hydrogen to uranium-235 atom concentration for an energy-independent extrapolation distance was successful in predicting the critical heights of various unreflected cylindrical assemblies. The critical dimensions of fully water-reflected cylindrical assemblies were reasonably well predicted using the theoretical buckling curve and reflector savings for equivalent spherical assemblies.

Fieno, Daniel↗

Control Architecture for Robotic Agent Command and Sensing

Control Architecture for Robotic Agent Command and Sensing (CARACaS) is a recent product of a continuing effort to develop architectures for controlling either a single autonomous robotic vehicle or multiple cooperating but otherwise autonomous robotic vehicles. CARACaS is potentially applicable to diverse robotic systems that could include aircraft, spacecraft, ground vehicles, surface water vessels, and/or underwater vessels. CARACaS incudes an integral combination of three coupled agents: a dynamic planning engine, a behavior engine, and a perception engine. The perception and dynamic planning en - gines are also coupled with a memory in the form of a world model. CARACaS is intended to satisfy the need for two major capabilities essential for proper functioning of an autonomous robotic system: a capability for deterministic reaction to unanticipated occurrences and a capability for re-planning in the face of changing goals, conditions, or resources. The behavior engine incorporates the multi-agent control architecture, called CAMPOUT, described in An Architecture for Controlling Multiple Robots (NPO-30345), NASA Tech Briefs, Vol. 28, No. 11 (November 2004), page 65. CAMPOUT is used to develop behavior-composition and -coordination mechanisms. Real-time process algebra operators are used to compose a behavior network for any given mission scenario. These operators afford a capability for producing a formally correct kernel of behaviors that guarantee predictable performance. By use of a method based on multi-objective decision theory (MODT), recommendations from multiple behaviors are combined to form a set of control actions that represents their consensus. In this approach, all behaviors contribute simultaneously to the control of the robotic system in a cooperative rather than a competitive manner. This approach guarantees a solution that is good enough with respect to resolution of complex, possibly conflicting goals within the constraints of the mission to be accomplished by the vehicle(s).

Huntsberger, Terrance↗

A Performance-Portable MultiGPU Implementation of 3D Euler Equations using ProtoX and IRIS

Computational scientists often face challenges when developing and optimizing code for high-performance computing (HPC), especially when trying to leverage GPUs. Given the heterogeneity of the nodes that comprise many modern HPC facilities, considerable demand exists for performance portable solutions for the core computational kernels used in many scientific computing libraries. In this work, we demonstrate a fourth-order finite volume method–based implementation of the Euler equations, which are an integral part of computational fluid dynamics. Our performance-portable multiGPU implementation for Euler equations uses ProtoX to generate kernels and IRIS for portability. ProtoX is a domain-specific language that uses a structured-grid partial differential equation library called Proto as its front end and the SPIRAL code generation system as its back end to generate optimized kernels for different architectures. Optimized kernels generated by ProtoX are orchestrated through the IRIS intelligent runtime system to provide portability. Two levels of optimizations within the IRIS runtime— directed acyclic graph fusion and task fusion—are explored to efficiently utilize computing resources in a multiGPU environment. Performance improvement through these optimizations is showcased by comparing the base ProtoX-IRIS implementation on AMD GPUs (Frontier node) and on NVIDIA GPUs (NVIDIA DGX-1).

Mankad, Het↗

Fast truncated SVD of sparse and dense matrices on graphics processors

We investigate the solution of low-rank matrix approximation problems using the truncated singular value decomposition (SVD). For this purpose, we develop and optimize graphics processing unit (GPU) implementations for the randomized SVD and a blocked variant of the Lanczos approach. Our work takes advantage of the fact that the two methods are composed of very similar linear algebra building blocks, which can be assembled using numerical kernels from existing high-performance linear algebra libraries. Furthermore, the experiments with several sparse matrices arising in representative real-world applications and synthetic dense test matrices reveal a performance advantage of the block Lanczos algorithm when targeting the same approximation accuracy.

Computer Science↗

A generalized Selberg zeta function for flat space cosmologies

Flat space cosmologies (FSCs) are time dependent solutions of three-dimensional (3D) gravity with a vanishing cosmological constant. They can be constructed from a discrete quotient of empty 3D flat spacetime and are also called shifted-boost orbifolds. Using this quotient structure, we build a new and generalized Selberg zeta function for FSCs, and show that it is directly related to the scalar 1-loop partition function. We then propose an extension of this formalism applicable to more general quotient manifolds $\mathcal{M}$/ℤ, based on representation theory of fields propagating on this background. Our prescription constitutes a novel and expedient method for calculating regularized 1-loop determinants, without resorting to the heat kernel. We compute quasinormal modes in the FSC using the zeroes of a Selberg zeta function, and match them to known results.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Sparse-Stochastic Fragmented Exchange for Large-Scale Hybrid Time-Dependent Density Functional Theory Calculations

Here we extend our recently developed sparse-stochastic fragmented exchange formalism for ground-state near-gap hybrid DFT to calculate absorption spectra within linear-response time-dependent generalized Kohn-Sham DFT (LR-GKS-TDDFT) for systems consisting of thousands of valence electrons within a grid-based/plane-wave representation. A mixed deterministic/fragmented-stochastic compression of the exchange kernel, here using long-range explicit exchange functionals, provides an efficient method for accurate optical spectra. Both real-time propagation as well as frequency-resolved Casida-equation-type approaches for spectra are presented, and the method is applied to large molecular dyes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Numerical solution of random singular integral equation appearing in crack problems

The solution of several elasticity problems, and particularly crack problems, can be reduced to the solution of one-dimensional singular integral equations with a Cauchy-type kernel or to a system of uncoupled singular integral equations. Here a method for the numerical solution of random singular integral equations of Cauchy type is presented. The solution technique involves a Chebyshev series approximation, the coefficients of which are the solutions of a system of random linear equations. This method is applied to the problem of periodic array of straight cracks inside an infinite isotropic elastic medium and subjected to a nonuniform pressure distribution along the crack edges. The statistical properties of the random solution are evaluated numerically, and the random solution is used to determine the values of the stress-intensity factors at the crack tips. The error, expressed as the difference between the mean of the random solution and the deterministic solution, is established. Values of stress-intensity factors at the crack tip for different random input functions are presented.

Sambandham, M.↗

Forward variable selection enables fast and accurate dynamic system identification with Karhunen-Loève decomposed Gaussian processes

A promising approach for scalable Gaussian processes (GPs) is the Karhunen-Loève (KL) decomposition, in which the GP kernel is represented by a set of basis functions which are the eigenfunctions of the kernel operator. Such decomposed kernels have the potential to be very fast, and do not depend on the selection of a reduced set of inducing points. However KL decompositions lead to high dimensionality, and variable selection thus becomes paramount. This paper reports a new method of forward variable selection, enabled by the ordered nature of the basis functions in the KL expansion of the Bayesian Smoothing Spline ANOVA kernel (BSS-ANOVA), coupled with fast Gibbs sampling in a fully Bayesian approach. It quickly and effectively limits the number of terms, yielding a method with competitive accuracies, training and inference times for tabular datasets of low feature set dimensionality. Theoretical computational complexities are O ( N P 2 ) in training and O ( P ) per point in inference, where N is the number of instances and P the number of expansion terms. The inference speed and accuracy makes the method especially useful for dynamic systems identification, by modeling the dynamics in the tangent space as a static problem, then integrating the learned dynamics using a high-order scheme. The methods are demonstrated on two dynamic datasets: a ‘Susceptible, Infected, Recovered’ (SIR) toy problem, along with the experimental ‘Cascaded Tanks’ benchmark dataset. Comparisons on the static prediction of time derivatives are made with a random forest (RF), a residual neural network (ResNet), and the Orthogonal Additive Kernel (OAK) inducing points scalable GP, while for the timeseries prediction comparisons are made with LSTM and GRU recurrent neural networks (RNNs) along with the SINDy package.

Hayes, Kyle↗

A two-level GPU-accelerated incomplete LU preconditioner for general sparse linear systems

This paper presents a parallel preconditioning approach based on incomplete LU (ILU) factorizations in the framework of Domain Decomposition (DD) for general sparse linear systems. We focus on distributed memory parallel architectures, specifically, those that are equipped with graphic processing units (GPUs). In addition to block-Jacobi, we present general purpose two-level ILU Schur complement-based approaches, where different strategies are presented to solve the coarse-level reduced system. These strategies are combined with modified ILU methods in the construction of the coarse-level operator, in order to effectively remove smooth errors by targeting an algebraically smooth vector. We leverage available GPU-based sparse matrix kernels to accelerate the setup and the solve phases of the proposed ILU preconditioner. We evaluate the efficiency of the proposed methods as a smoother for algebraic multigrid (AMG) and as a preconditioner for Krylov subspace methods on challenging anisotropic diffusion problems and a collection of general sparse matrices.

97 MATHEMATICS AND COMPUTING↗

Performance Optimization Methods for a Memory-Bound, Unstructured-Grid CFD Application on Massively Parallel GPU Platforms

Computational performance of the FUN3D unstructured-grid computational fluid dynamics (CFD) application on massively parallel GPU environments is memory-bound and highly dependent upon efficient reads from and atomic updates to the irregular cell-, edge-, and node-based data structures. In this talk, we present recent efforts into optimizing select performance-critical kernels on NVIDIA Tesla V100 and A100 GPUs and AMD CDNA MI100 GPUs. A novel use of L2 cache residency controls and asynchronous loads into on-chip shared memory are explored on the A100 GPU for the sparse iterative solver, which is dominated by mixed-precision, sparse matrix vector multiplication. Demonstrations show that these methods improve global memory bandwidth utilization by 13.5% on the A100 GPU. Several techniques are also presented that use registers and/or shared memory to facilitate array transposition and aggregation which combine to reduce the frequency and increase the cache efficiency of floating-point atomic updates to the irregular data structures. These methods are demonstrated to improve the kernel throughput by nearly 500% on select kernels on the AMD MI100 over atomic updates directly to global memory. Overall, both V100 and A100 GPUs outperformed the MI100 GPU on kernels dominated by double-precision atomic updates; however, the techniques demonstrated here reduced the performance gap and improved the MI100 performance.

GPU CPU unstructured CFD memory↗

Explosive Soot Challenge (Final Report)

This project assembled a broad ensemble of modeling and experimentation tools to study the morphological and optical properties of detonation soots in explosive fireballs. A gram-scale hemispherical high explosive was studied in a low-pressure controlled environment using in-situ experimentation with diffusely illuminated visible absorption spectroscopy, particle sizing through light scattering techniques, and post-test collections with subsequent morphological analysis. Hydrocode modeling was performed to replicate the detonation flow observations, and subsequent aerosol kinetics models provided particle size distributions and extinction coefficients from the hydrocode results. Experimentally observed soot morphologies agreed with expectation from the literature - a bimodal distribution was found, brought upon by the particles growing to a size where their inertia and fluid wakes are non-negligible. The aerosol kinetics model did not replicate the observed bimodal size distribution for lack of a coagulation kernel to represent the behavior. To recover particulate optical properties, a spectrally resolved absorption spectroscopy method termed Spectral diffuse back-illuminated extinction imaging (SBI-EI) was developed and implemented on two explosive types. Inverting the absorption spectra using a Kramers-Kronig consistent method yielded the complex index of refraction for the soots produced by the explosives. This method resulted in an unrealistic index of refraction for one of the two explosives, and this is suspected to be due to the model neglecting scattering brought upon by the large particle sizes observed. In addition to the core work, three additional studies were performed in parallel. These investigated the impact of scattering on diffuse absorption spectroscopy, studied how soots oxidate and sublimate in a well-controlled shock tube, and laid the theoretical groundwork for a new collision kernel to replicate the bimodal size distribution from the observations. Summaries of these efforts are included at the end of this report.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Scalable Risk Assessment of Rare Events in Power Systems With Uncertain Wind Generation and Loads

Risk assessment of rare events has become increasingly important in power system planning and operation with the increasing integration of renewable energy and the presence of system uncertainties. However, quantifying the risk posed by rare events via the traditional method, i.e., Monte Carlo sampling (MCS), incurs substantial computational expense stemming from the vast ensemble of power flow simulations. To accelerate the assessment, this paper proposes a Deep Neural Network (DNN)-kernelized vector-valued Gaussian Process (VVGP) approach with excellent computational efficiency while maintaining high accuracy. Consequently, serving as a surrogate model for the power flow solver, the DNN-kernelized VVGP enables significantly faster but accurate risk assessment compared to the power flow solver. The developed surrogate model evaluates low-order N - k events that contain more than 90% instances by adeptly capturing the topological features while the high-order N - k events are assessed via a power flow solver, thereby striking a balance between computational efficiency and uncertainty quantification accuracy. Moreover, the model incorporates a Support Vector Machine (SVM) classifier to resample concerning low-probability tail events to counteract the biases potentially introduced during the DNN-kernelized VVGP evaluations. Simulations conducted on the modified IEEE 24-bus, 118-bus, and European 1354-bus systems demonstrate that the proposed method maintains the accuracy benchmark set by MCS while significantly reducing computational demands in large-scale power systems as compared to other state-of-the-art methods.

17 WIND ENERGY↗

The effect of atmospheric transmissivity on model and observational estimates of the sea ice albedo feedback

The sea-ice-albedo feedback (SIAF) is the product of the ice sensitivity (IS) – how much the surface albedo in sea-ice regions changes as the planet warms– and the radiative sensitivity (RS) – how much the top of atmosphere radiation changes as the surface albedo changes. We demonstrate that the RS calculated from radiative kernels in climate models is reproduced from calculations using the “approximate partial radiative perturbation” method that uses the climatological radiative fluxes at the top of atmosphere and the assumption that the atmosphere is isotropic to shortwave radiation. This method facilitates the comparison of RS from satellite-based estimates of climatological radiative fluxes with RS estimates across a full suite of coupled climate models and, thus, allows model evaluation of a quantity important in characterizing the climate impact of sea ice concentration changes. The satellite based RS is within the model range of RS that differs by a factor of two across climate models in both the Arctic and Southern Ocean. Observed trends in Arctic sea ice are used to estimate IS which, in conjunction with the satellite-based RS yields an SIAF of 0.16 ± 0.04 W m-2 K-1. This Arctic SIAF estimate suggests a modest amplification of future global surface temperature change by approximately 14% relative to a climate system with no SIAF. We calculate the global albedo feedback in climate models using model specific RS and IS and find a model mean feedback parameter of 0.37 W m-2 K-1 which is 40% larger than the IPCC AR5 estimate based on using RS calculated from radiative kernel calculations in a single climate model.

Arctic, Antarctic, climate change, Climate & Earth↗