Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “kernel method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35

Real-space inversion and super-resolution of ultrafast scattering

Ultrafast scattering using x-rays or electrons is an emerging method to obtain structure dynamics at the atomic lengthscales and timescales. However, directly resolving in real-space atomic motions is inherently limited by the finite detector range and the probe energy. So, as a result, the time-resolved signal interpretation is mostly done in reciprocal space and relies on modeling and simulations of specific structures and processes. Here, we introduce a model-free approach to directly resolve scattering signals in real space, surpassing the diffraction limit, using scattering kernels and signal priors that naturally arise from the measurement constraints. We demonstrate the approach on simulated and experimental data, recover multiple atomic motions at sub-angstrom resolutions, and discuss the recovery accuracy and resolution limits versus signal fidelity. The approach offers a robust path to obtain high-resolution real-space information of atomic-scale structure dynamics using current time-resolved x-ray or electron scattering sources.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Protein Kinase Classification with 2866 Hidden Markov Models and One Support Vector Machine

The main application considered in this paper is predicting true kinases from randomly permuted kinases that share the same length and amino acid distributions as the true kinases. Numerous methods already exist for this classification task, such as HMMs, motif-matchers, and sequence comparison algorithms. We build on some of these efforts by creating a vector from the output of thousands of structurally based HMMs, created offline with Pfam-A seed alignments using SAM-T99, which then must be combined into an overall classification for the protein. Then we use a Support Vector Machine for classifying this large ensemble Pfam-Vector, with a polynomial and chisquared kernel. In particular, the chi-squared kernel SVM performs better than the HMMs and better than the BLAST pairwise comparisons, when predicting true from false kinases in some respects, but no one algorithm is best for all purposes or in all instances so we consider the particular strengths and weaknesses of each.

Weber, Ryan↗

Next-Cycle Optimal Fuel Control for Cycle-to-Cycle Variability Reduction in EGR-Diluted Combustion

In this simulation study, cycle-to-cycle fuel control was used to reduce CCV by injecting additional fuel in operating conditions with sporadic misfires and partial burns. An optimal control policy was proposed that utilizes 1) a physics-based model that tracks in-cylinder gas composition and 2) a one-step-ahead prediction of the combustion efficiency based on a kernel density estimator. The optimal solution, however, presents a tradeoff between the reduction in combustion CCV and the increase in fuel injection quantity required to stabilize the charge. Such a tradeoff can be ad- just by a single parameter embedded in the cost function.

Maldonado, BryanP. [Oak Ridge National Lab. (ORNL)↗

The Kokkos Ecosystem [Brief]

In 2016/2017, the field of High-Performance Computing (HPC) entered a new era driven by fundamental physics challenges to produce ever more energy and cost-efficient processors. Since the convergence on the Message-Passing Interface (MPI) standard in the mid-1990s, application developers enjoyed a seemingly static view of the underlying machine — that of a distributed collection of homogeneous nodes executing in collaboration. However, after almost two decades of dominance, the sole use of MPI to derive parallelism acted as a limiter to improved future performance. While MPI is widely expected to continue to function as the basic mechanism for communication between compute nodes for the immediate future, additional parallelism is required on the computing node itself if high performance and efficiency goals are to be realized. When reviewing the architectures of the top HPC systems today, the change in paradigm is clear: the compute nodes of the leading machines in the world are either powered by many-core chips with a few dozen cores each, or use heterogeneous designs, where traditional CPUs marshal work to massively parallel compute accelerators which has as many as 200,000 processing threads in flight simultaneously. Complicating matters further for application developers, each processor vendor has its own preferred way of writing code for their architecture.The Kokkos EcoSystem was released by Sandia in 2017 to address this new era in HPC system design by providing a vendor independent performance portable programming system for scientific, engineering, and mathematical software applications written in the C++ programming language. Using Kokkos, application developers can be more productive because they will not have to create and maintain separate versions of their software for each architecture, nor will they have to be experts in each architecture's peculiar requirements. Instead, they will have a single method of programming for the diverse set of modern HPC architectures. While Kokkos started in 2011 as a programming model only, it soon became clear that complex applications needed more. It is also critical to have a portable mathematical functions and developers need tools to debug their applications, gain insight into the performance characteristics of their codes and tune algorithm performance parameters through automated processes. The Kokkos EcoSystem addresses those needs through its three main components: the Kokkos Core programming model, the Kokkos Kernels math library, and the Kokkos Tools project.

97 MATHEMATICS AND COMPUTING↗

Mixed boundary-value problems in mechanics

Definitions in the case of multiple series equations and multiple integral equations are examined. In considering the solution of a given mixed boundary value problem perhaps the simplest technique is the direct application of the method of complex potentials provided the problem admits such potentials and the domain and the boundary conditions are suitable for such an application. The direct application of complex potentials is described with the aid of examples, taking into account a problem in potential theory, the case of periodic cuts, and an elasticity problem for a nonhomogeneous plane. The reduction to singular integral equations is discussed along with the numerical solution of singular integral equations of the first kind, integral equations with generalized Cauchy kernels, and singular integral equations of the second kind.

Erdogan, F.↗

High-performance strategies for the recent MRSF-TDDFT in GAMESS

Multiple ERI (Electron Repulsion Integral) tensor contractions (METC) with several matrices are ubiquitous in quantum chemistry. In response theories, the contraction operation, rather than ERI computations, can be the major bottleneck, as its computational demands are proportional to the multiplicatively combined contributions of the number of excited states and the kernel pre-factors. Here, this paper presents several high-performance strategies for METC. Optimal approaches involve either the data layout reformations of interim density and Fock matrices, the introduction of intermediate ERI quartet buffer, and loop-reordering optimization for a higher cache hit rate. The combined strategies remarkably improve the performance of the MRSF (mixed reference spin flip)-TDDFT (time-dependent density functional theory) by nearly 300%. The results of this study are not limited to the MRSF-TDDFT method and can be applied to other METC scenarios.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Efficient quadrature rules for finite element discretizations of nonlocal equations

In this paper we design efficient quadrature rules for finite element discretizations of nonlocal diffusion problems with compactly supported kernel functions. Two of the main challenges in nonlocal modeling and simulations are the prohibitive computational cost and the nontrivial implementation of discretization schemes, especially in three-dimensional settings. In this work we circumvent both challenges by introducing a parametrized mollifying function that improves the regularity of the integrand, utilizing an adaptive integration technique, and exploiting parallelization. We first showthat the “mollified” solution converges to the exact one as the mollifying parameter vanishes, then we illustrate the consistency and accuracy of the proposed method on several two- and three-dimensional test cases. Furthermore, we demonstrate the good scaling properties of the parallel implementation of the adaptive algorithm and we compare the proposed method with recently developed techniques for efficient finite element assembly.

97 MATHEMATICS AND COMPUTING↗

CSPlib - A Software Toolkit for the Analysis of Dynamical Systems and Chemical Kinetic Models

CSPlib is an open source software library for analyzing general ordinary differential equation (ODE) systems and detailed chemical kinetic ODE systems. It relies on the computational singular perturbation (CSP) method for the analysis of these systems. The software provides support for: General ODE models (gODE model class) for computing source terms and Jacobians for a generic ODE system; TChem model (ChemElemODETChem model class) for computing source term, Jacobian, other necessary chemical reaction data, as well as the rates of progress for a homogenous batch reactor using an elementary step detailed chemical kinetic reaction mechanism. This class relies on the TChem [2] library; A set of functions to compute essential elements of CSP analysis (Kernel class). This includes computations of the eigensolution of the Jacobian matrix, CSP basis vectors and co-vectors, time scales (reciprocals of the magnitudes of the Jacobian eigenvalues), mode amplitudes, CSP pointers, and the number of exhausted modes. This class relies on the Tines library; A set of functions to compute the eigensolution of the Jacobian matrix using Tines library GPU eigensolver; A set of functions to compute CSP indices (Index Class). This includes participation indices and both slow and fast importance indices.

97 MATHEMATICS AND COMPUTING↗

Efficient quadrature rules for finite element discretizations of nonlocal equations

Here, in this paper, we design efficient quadrature rules for finite element (FE) discretizations of nonlocal diffusion problems with compactly supported kernel functions. Two of the main challenges in nonlocal modeling and simulations are the prohibitive computational cost and the nontrivial implementation of discretization schemes, especially in three-dimensional settings. In this work, we circumvent both challenges by introducing a parametrized mollifying function that improves the regularity of the integrand, utilizing an adaptive integration technique, and exploiting parallelization. We first show that the “mollified” solution converges to the exact one as the mollifying parameter vanishes, then we illustrate the consistency and accuracy of the proposed method on several two- and three-dimensional test cases. Furthermore, we demonstrate the good scaling properties of the parallel implementation of the adaptive algorithm and we compare the proposed method with recently developed techniques for efficient FE assembly.

97 MATHEMATICS AND COMPUTING↗

Revisiting Temporal Blocking Stencil Optimizations

Iterative stencils are used widely across the spectrum of High Performance Computing (HPC) applications. Many efforts have been put into optimizing stencil GPU kernels, given the prevalence of GPU-accelerated supercomputers. To improve the data locality, temporal blocking is an optimization that combines a batch of time steps to process them together. Under the observation that GPUs are evolving to resemble CPUs in some aspects, we revisit temporal blocking optimizations for GPUs. We explore how temporal blocking schemes can be adapted to the new features in the recent Nvidia GPUs, including large scratchpad memory, hardware prefetching, and device-wide synchronization. We propose a novel temporal blocking method, EBISU, which champions low device occupancy to drive aggressive deep temporal blocking on large tiles that are executed tile-by-tile. We compare EBISU with state-of-the-art temporal blocking libraries: STENCILGEN and AN5D. We also compare with state-of-the-art stencil auto-tuning tools that are equipped with temporal blocking optimizations: ARTEMIS and DRSTENCIL. Over a wide range of stencil benchmarks, EBISU achieves speedups up to 2.53x and a geometric mean speedup of 1.49x over the best state-of-the-art performance in each stencil benchmark.

Zhang, Lingqi↗

Scalable quantum processor noise characterization

Measurement fidelity matrices (MFMs) (also called error kernels) are a natural way to characterize state preparation and measurement errors in near-term quantum hardware. They can be employed in post processing to mitigate errors and substantially increase the effective accuracy of quantum hardware. However, the feasibility of using MFMs is currently limited as the experimental cost of determining the MFM for a device grows exponentially with the number of qubits. In this work we present a scalable way to construct approximate MFMs for many-qubit devices based on cumulant expansions. Our method can also be used to characterize various types of correlation error.

Hamilton, Kathleen↗

Atomistic and mesoscale simulations to determine effective diffusion coefficient of fission products in SiC

The silicon carbide (SiC) layer in tristructural isotropic (TRISO) particles serves as the barrier to prevent escape of fission products produced in the fuel kernel. Knowing the diffusion coefficient of fission products through SiC is critical to determining whether fission gas can escape from the particle. It has been observed in experiments that Ag accumulated in grain boundaries and triple junctions in SiC. It is hypothesized that grain boundary diffusion is the primary pathway by which fission products penetrate the SiC layer. In this report, the effective diffusion coefficient of the fission product Ag through the grain boundary network is calculated using a combination of atomistic and phase-field methods. The grain boundary diffusion coefficient is calculated using molecular dynamics simulations. The bulk diffusion coefficient is determined using a combination of density functional theory and nudged elastic band methods. An effective diffusion coefficient is calculated, accounting for the grain structure using a phase-field method. The effective diffusion coefficient will be incorporated into Bison and fission product release calculations are compared to available experimental data.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

GPU Profiling and Optimizing xRAGE (Final Report)

Our project’s objective is to increase the efficiency of GPU-enabled kernels in xRAGE. To do so, we conduct GPU profiling with NSight Systems on xRAGE tests unsplit_sod_1d and unsplit_sedov_2d to identify bottlenecks and understand the behavior of the GPU during code execution. Next, we analyze these generated GPU profiles to locate the lines of code whose optimization have the most potential for improving runtime. We replicate the structure of the code in smaller test problems that are easier to understand, edit, and run quickly. Within these test problems, we implement two different methods of improving performance: transformation of nested loops into a single MDRangePolicy and hierarchical parallelization using teams of threads. Both methods show speedups in the test code, and after transferring them to xRAGE, they both show up to 30x speedups on various computing platforms. Profiling the edited versions of xRAGE reveals that the GPU successfully executed the bottlenecks with greater efficiency

97 MATHEMATICS AND COMPUTING↗

Variable Order and Distributed Order Fractional Operators

Many physical processes appear to exhibit fractional order behavior that may vary with time or space. The continuum of order in the fractional calculus allows the order of the fractional operator to be considered as a variable. This paper develops the concept of variable and distributed order fractional operators. Definitions based on the Riemann-Liouville definitions are introduced and behavior of the operators is studied. Several time domain definitions that assign different arguments to the order q in the Riemann-Liouville definition are introduced. For each of these definitions various characteristics are determined. These include: time invariance of the operator, operator initialization, physical realization, linearity, operational transforms. and memory characteristics of the defining kernels. A measure (m2) for memory retentiveness of the order history is introduced. A generalized linear argument for the order q allows the concept of "tailored" variable order fractional operators whose a, memory may be chosen for a particular application. Memory retentiveness (m2) and order dynamic behavior are investigated and applications are shown. The concept of distributed order operators where the order of the time based operator depends on an additional independent (spatial) variable is also forwarded. Several definitions and their Laplace transforms are developed, analysis methods with these operators are demonstrated, and examples shown. Finally operators of multivariable and distributed order are defined in their various applications are outlined.

Lorenzo, Carl F.↗

Enhanced relaxed physical factorization preconditioner for coupled poromechanics

The relaxed physical factorization (RPF) preconditioner is a recent algorithm allowing for the efficient and robust solution to the block linear systems arising from the three-field displacement-velocity-pressure formulation of coupled poromechanics. For its application, however, it is necessary to invert blocks with the algebraic form C^ = (C + βFF T ), where C is a symmetric positive definite matrix, FF T a rank-deficient term, and β a real non-negative coefficient. The inversion of C^, performed in an inexact way, can become unstable for large values of β, as it usually occurs at some stages of a full poromechanical simulation. In this work, we propose a family of algebraic techniques to stabilize the inexact solve with C^. This strategy can prove useful in other problems as well where such an issue might arise, such as augmented Lagrangian preconditioning techniques for Navier-Stokes or incompressible elasticity. First, we introduce an iterative scheme obtained by a natural splitting of matrix C^. Second, we develop a technique based on the use of a proper projection operator annihilating the near-kernel modes of C^. Both approaches give rise to a novel class of preconditioners denoted as Enhanced RPF (ERPF). Furthermore, effectiveness and robustness of the proposed algorithms are demonstrated in both theoretical benchmarks and real-world large-size applications, outperforming the native RPF preconditioner.

97 MATHEMATICS AND COMPUTING↗

A Probabilistic Scheme for Semilinear Nonlocal Diffusion Equations with Volume Constraints

This work presents a probabilistic scheme for solving semilinear nonlocal diffusion equations with volume constraints and integrable kernels. The nonlocal model of interest is defined by a time-dependent semilinear partial integro-differential equation (PIDE), in which the integro-differential operator consists of both local convection-diffusion and nonlocal diffusion operators. Here, our numerical scheme is based on the direct approximation of the nonlinear Feynman–Kac formula that establishes a link between nonlinear PIDEs and stochastic differential equations. The exploitation of the Feynman–Kac representation avoids solving dense linear systems arising from nonlocal operators. Compared with existing stochastic approaches, our method can achieve first-order convergence after balancing the temporal and spatial discretization errors, which is a significant improvement of existing probabilistic/stochastic methods for nonlocal diffusion problems. Error analysis of our numerical scheme is established. The effectiveness of our approach is shown in two numerical examples. The first example considers a three-dimensional nonlocal diffusion equation to numerically verify the error analysis results. The second example presents a physics problem motivated by the study of heat transport in magnetically confined fusion plasmas.

97 MATHEMATICS AND COMPUTING↗

A dual reciprocal boundary element formulation for viscous flows

The advantages inherent in the boundary element method (BEM) for potential flows are exploited to solve viscous flow problems. The trick is the introduction of a so-called dual reciprocal technique in which the convective terms are represented by a global function whose unknown coefficients are determined by collocation. The approach, which is necessarily iterative, converts the governing partial differential equations into integral equations via the distribution of fictitious sources or dipoles of unknown strength on the boundary. These integral equations consist of two parts. The first is a boundary integral term, whose kernel is the unknown strength of the fictitious sources and the fundamental solution of a convection-free flow problem. The second part is a domain integral term whose kernel is the convective portion of the governing PDEs. The domain integration can be transformed to the boundary by using the dual reciprocal (DR) concept. The resulting formulation is a pure boundary integral computational process.

Lafe, Olu↗

Optimized Umkehr Profile Algorithm for Ozone Trend Analyses

The long-term record of Umkehr measurements from four NOAA Dobson spectrophotometers was reprocessed after updates to the instrument calibration procedures. In addition, a new data quality-control tool was developed for the Dobson automation software (WinDobson). This paper presents a comparison of Dobson Umkehr ozone profiles from NOAA ozone network stations (Boulder, OHP, MLO, Lauder) against several satellite records, including Aura Microwave Limb Sounder (MLS; ver. 4.2), and combined SBUV and OMPS records (NASA AGG and NOAA COH). A subset of satellite data is selected to match Dobson Umkehr observations at each station spatially (distance less than 200 km) and temporally (within 24 hours). Umkehr Averaging Kernels (AKs) are applied to vertically smooth all overpass satellite profiles prior to comparisons. The station Umkehr record consists of several instrumental records, which have different optical characterizations, and thus instrument-specific stray light contributes to the data processing errors and creates step changes in the record. This work evaluates the overall quality of Umkehr long-term measurements at NOAA ground-based stations and assesses the impact of the instrumental changes on the stability of the Umkehr ozone profile record. This paper describes a method designed to correct biases and discontinuities in the retrieved Umkehr profile that originate from the Dobson calibration process, repair, or optical realignment of the instrument. The M2GMI and GMI CTM ozone profile model output matched to station location and date of observation is used to evaluate instrumental step changes in the Umkehr record. Homogenization of the Umkehr record and discussion of the apparent stray light error in retrieved ozone profiles are the focus of this paper. Homogenization of ground-based records is of great importance for studies of long-term ozone trends and climate change.

Umkehr↗