Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “kernel method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

Spectrally Stabilized Interface Capturing Formulation and Implementation in Nek5000/NekRS

This report documents the formulation of a novel level-set method for incompressible two-phase flows in the continuous Galerkin (CG) high order spectral element framework. The overall method hinges on a novel implementation of the spectral vanishing viscosity (SVV) operator for the stabilization of linear/non-linear hyperbolic problems. The multidimensional SVV convolution kernels, which in essence, have a similar effect as a high pass filter applied to the derivatives, are formulated by exploiting the tensor product form, analogous to the construction of the usual stiffness matrix system. The resulting kernels are directionally decoupled and ensure a linear, symmetric positive definite, elliptic matrix operator. The SVV formulation is demonstrated to provide a robust stabilizing mechanism through challenging linear and non-linear hyperbolic problems, including problems pertinent to the level-set formulation. The two-phase framework conceptualized herein is based on the conservative level-set (CLS) method which represents the interface between the fluids by the 0.5 iso-contour of the smoothed Heaviside function. The CLS method is augmented with a preconditioning procedure for interface normals using the signed distance function which precludes the manifestation of spurious oscillations in the vicinty of the interface. Further, the existing mixed explicit-implicit approach for the solution of Navier-Stokes equations in Nek5000, as described in Tomboulides et al, is augmented with a pressure coefficient splitting approach for the Poisson equation, which greatly accelerated the convergence of pressure solver for two-phase systems with large density ratio. The robustness and accuracy of the overall two-phase method is demonstrated through canonical challenging problems involving high density and viscosity ratios, with and without surface tension. The two-phase formulation is wholly implemented in Nek5000 and the SVV stabilization method is implemented in NekRS, which is the essential precursor to the two-phase framework, undergoing active development.

97 MATHEMATICS AND COMPUTING↗

Iterative Discrete Ordinates Solution of the Equation for the Surface-Reflected Radiance

This paper presents a new method of numerical solution of the integral equation for the radiance reflected from an anisotropic surface. The equation relates the radiance at the surface level with BRDF and solutions of the standard radiative transfer problems for a slab with no reflection on its surfaces. It is also shown that the kernel of the equation satisfies the condition of the existence of a unique solution and the convergence of the successive approximations to that solution. The developed method features two basic steps: discretization on a 2D quadrature, and solving the resulting system of algebraic equations with successive over-relaxation method based on the Gauss-Seidel iterative process. Presented numerical examples show good coincidence between the surface-reflected radiance obtained with DISORT and the proposed method. Analysis of contributions of the direct and diffuse (but not yet reflected) parts of the downward radiance to the total solution is performed. Together, they represent a very good initial guess for the iterative process. This fact ensures fast convergence. The numerical evidence is given that the fastest convergence occurs with the relaxation parameter of 1 (no relaxation). An integral equation for BRDF is derived as inversion of the original equation. The potential of this new equation for BRDF retrievals is analyzed. The approach is found not viable as the BRDF equation appears to be an ill-posed problem, and it requires knowledge the surface-reflected radiance on the entire domain of both Sun and viewing zenith angles.

Alexander Radkevich↗

Lagrangian–Eulerian multidensity topology optimization with the material point method

Abstract In this paper, a hybrid Lagrangian–Eulerian topology optimization (LETO) method is proposed to solve the elastic force equilibrium with the Material Point Method (MPM). LETO transfers density information from freely movable Lagrangian carrier particles to a fixed set of Eulerian quadrature points. This transfer is based on a smooth radial kernel involved in the compliance objective to avoid the artificial checkerboard pattern. The quadrature points act as MPM particles embedded in a lower‐resolution grid and enable a subcell multidensity resolution of intricate structures with a reduced computational cost. A quadrature‐level connectivity graph‐based method is adopted to avoid the artificial checkerboard issues commonly existing in multiresolution topology optimization methods. Numerical experiments are provided to demonstrate the efficacy of the proposed approach.

Li, Yue↗

Evaluation of XCT for Matrix Density Measurement of Particle Fuel Forms

Particle fuel forms generally consist of a dispersion of fuel, such as tristructural isotropic (TRISO) particles, within a refractory matrix (e.g., graphite or silicon carbide). The density of matrix materials for particle fuel forms is of interest for modeling fuel form strength and thermal properties and may be specified as a quality control parameter, depending on reactor design. Some of the uncertainty associated with traditional, manual approaches can be eliminated by performing x-ray computed tomography (XCT) on the fuel forms and applying image processing methods to generate a precise count of the number of particles. This also removes the need to include determination of particle count within each individual fuel form during fabrication. Unfortunately, reconstruction artifacts from high-Z uranium-bearing kernels prevent accurate measurement of individual particle volumes using this approach, so the use of mean particle mass and volume are still necessary for computation of average fuel form matrix density. This method of using XCT to count particles in individual fuel form for determination of average matrix density was applied to three archived compacts from the Advanced Gas Reactor Fuel Development and Qualification (AGR)-1 campaign, four archived compacts with uranium carbide/uranium oxide (UCO) TRISO from the AGR-2 campaign, and three archived UO 2 -TRISO compacts from the AGR-2 campaign. The resulting density values were compared with those previously reported, showing slight changes due to uncertainties in the previously used number of particles in each of these cylindrical, graphite matrix compacts.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Spatial Signatures of Electron Correlation in Least-Squares Tensor Hypercontraction

Least Squares Tensor Hypercontraction (LS-THC) has received some attention in recent years as an approach to reduce the significant computational costs of wavefunc- tion based methods in quantum chemistry. However, previous work has demonstrated that the LS-THC factorization performs disproportionately worse in the description of wavefunction components (e.g. cluster amplitudes T 2 ) than Hamiltonian compo- nents (e.g. electron repulsion integrals (pq|rs)). This work develops novel theoretical methods to study the source of these errors in the context of the real-space T 2 kernel, and reports, for the first time, the existence of a “correlation feature” in the errors of the LS-THC representation of the “exchange-like” correlation energy EX and T 2 that is remarkably consistent across ten molecular species, three correlated wavefunctions, and four basis sets. This correlation feature portends the existence of a “pair-point kernel” missing in the usual LS-THC representation of the wavefunction, which critically depends upon pairs of grid points situated close to atoms and with inter-pair distances between one and two Bohr radii. These findings point the way for future LS-THC developments to address these shortcomings.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Simulation of electron Bernstein waves using FullWave with a 2D non-local hot plasma model

Hot plasma wave simulation capability is expanded in the FullWave code by updating the hybrid iterative solver in the code with a semi-implicit time stepping method. The new approach is used to simulate Electron Bernstein Wave (EBW) heating in over-dense spherical tokamak plasmas. The code’s hybrid iterative solver circumvents the prohibitive memory cost of direct methods by combining a time evolution of Maxwell’s equations with frequency-domain relaxation, while the conductivity kernel, calculated via 3D particle tracking, captures the essential non-local wave–particle interactions. One-dimensional EBW simulations verify the algorithm’s accuracy by demonstrating mode conversion from X-mode wave to EBW at the upper hybrid resonance and a strong cyclotron damping near the plasma core. Two-dimensional simulation reproduces the predicted short EBW wavelength and quantitatively matches the hot-plasma dispersion relation. This study demonstrates the fidelity of the hybrid solver for the electron cyclotron frequency range.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Atwood effects on nonlocality of the scalar transport closure in Rayleigh-Taylor mixing

The importance of nonlocality is assessed in modeling mean scalar transport for turbulent Rayleigh-Taylor (RT) mixing at different Atwood numbers. Building on the two-dimensional incompressible work of Lavacot et al. [J. Fluid Mech. 985, A47 (2024)], the present work extends the macroscopic forcing method to variable density problems in three-dimensional space to measure moments of the generalized eddy diffusivity kernel in RT mixing for increasing Atwood numbers (𝐴 = 0.05, 0.3, 0.5, 0.8). It is found that as 𝐴 increases, (1) the eddy diffusivity moments become asymmetric and (2) the higher-order eddy diffusivity moments become larger relative to the leading-order diffusivity, indicating that nonlocality becomes more important at higher 𝐴. There is a particularly strong temporal nonlocality at higher 𝐴, suggesting stronger history effects. In conclusion, the implications of these findings for closure modeling for finite-Atwood RT are discussed.

general physics↗

LOGAN: High-Performance GPU-Based X-Drop Long-Read Alignment

Pairwise sequence alignment is one of the most computationally intensive kernels in genomic data analysis, accounting for more than 90% of the runtime for key bioinformatics applications. This method is particularly expensive for third-generation sequences due to the high computational cost of analyzing sequences of length between 1Kb and 1Mb. Given the quadratic overhead of exact pairwise algorithms for long alignments, the community primarily relies on approximate algorithms that search only for high-quality alignments and stop early when one is not found. In this work, we present the first GPU optimization of the popular X-drop alignment algorithm, that we named LOGAN. Results show that our high-performance multi-GPU implementation achieves up to 181.6 GCUPS and speed-ups up to 6.6× and 30.7× using 1 and 6 NVIDIA Tesla V100, respectively, over the state-of-the-art software running on two IBM Power9 processors using 168 CPU threads, with equivalent accuracy. We also demonstrate a 2.3× LOGAN speed-up versus ksw2, a state-of-art vectorized algorithm for sequence alignment implemented in minimap2, a long-read mapping software. Furthermore, to highlight the impact of our work on a real-world application, we couple LOGAN with a many-to-many long-read alignment software called BELLA, and demonstrate that our implementation improves the overall BELLA runtime by up to 10.6×. Finally, we adapt the Roofline model for LOGAN and demonstrate that our implementation is near optimal on the NVIDIA Tesla V100s.

97 MATHEMATICS AND COMPUTING↗

Numerical eigen-spectrum slicing, accurate orthogonal eigen-basis, and mixed-precision eigenvalue refinement using OpenMP data-dependent tasks and accelerator offload

Performing a variety of numerical computations efficiently and, at the same time, in a portable fashion requires both an overarching design followed by a number of implementation strategies. All of these are exemplified below as we present transitioning the PLASMA numerical library from relying on dependence-driven large tasks to achieving utilization of fine grain tasking and offload to hardware accelerators while keeping its core dependence sets: OpenMP source code pragmas and runtime for most system-level functionality and basic low-level numerical kernels provided directly by hardware vendors or open source projects with vendor contributions. We also present new algorithmic methods and their efficient parallel implementations including fine grained tasking for eigen-spectrum slicing and offload for mixed-precision eigenvalue refinement. We provide performance, scaling, and numerical results showing sizable gains over the available solutions from either the open source and vendor-provided packages.

Luszczek, Piotr↗

Retrieval of aerosol size distribution moments from multiwavelength particulate extinction measurements

Two methods for inferring aerosol size distribution moments from multiwavelength particulate extinction measurements are studied. The methods are an eigenvalue technique that approximates an appropriate moment-weighting function by a linear combination of kernel functions and a conversion ratio approach that uses the ratio of the particulate extinction measurements at two wavelengths to choose a model moment-to-extinction conversion ratio. The techniques are applied to infer the third moment, or volume, of the aerosol size distribution from actual particulate extinction measurements taken as part of the Stratospheric Aerosol and Gas Experiment II during a correlative measurement experiment in Brazil in April 1985.

Livingston, John M.↗

Enzymic synthesis of indole-3-acetyl-1-O-beta-d-glucose. I. Partial purification and characterization of the enzyme from Zea mays

The first enzyme-catalyzed reaction leading from indole-3-acetic acid (IAA) to the myo-inositol esters of IAA is the synthesis of indole-3-acetyl-1-O-beta-D-glucose from uridine-5'-diphosphoglucose (UDPG) and IAA. The reaction is catalyzed by the enzyme, UDPG-indol-3-ylacetyl glucosyl transferase (IAA-glucose-synthase). This work reports methods for the assay of the enzyme and for the extraction and partial purification of the enzyme from kernels of Zea mays sweet corn. The enzyme has an apparent molecular weight of 46,500 an isoelectric point of 5.5, and its pH optimum lies between 7.3 and 7.6. The enzyme is stable to storage at zero degrees but loses activity during column chromatographic procedures which can be restored only fractionally by addition of column eluates. The data suggest either multiple unknown cofactors or conformational changes leading to activity loss.

NASA Discipline Number 40-10↗

MOOSE ProbML: Parallelizable Probabilistic Machine Learning and Uncertainty Quantification Capabilities

The Multiphysics Object Oriented Simulation Environment (MOOSE) is a widely used open- source finite element software for performing multiphysics multiscale simulations in a massively parallel fashion. Recently, the computational team at Idaho National Laboratory (INL) has implemented Probabilistic Machine Learning (ProbML) capabilities in MOOSE—in a parallelized fashion—and enable active learning with large-scale computational models for tasks such as surrogate model development, scale bridging, forward/inverse uncertainty quantification (UQ), Bayesian optimization, etc. This presentation summarizes these developments in MOOSE along with demonstrations on several real applications relevant to nuclear energy. At the fundamental level, samplers like Monte Carlo/Latin Hypercube, variance reduction, parallelized Markov Chain Monte Carlo (MCMC) support uncertainty propagation in both forward and inverse settings. These samplers can be integrated with the Gaussian processes (GP) suite in MOOSE, which offer several variants like scalar GPs, multi-output GPs, and deep GPs, to enable active learning. These GPs can be tuned using gradient-based optimization methods like Adam and its variants or gradient-free methods like the elliptical slice sampler (a variant of MCMC adept under Gaussian settings) for more complex covariance kernels or likelihoods whose gradient computations can be cumbersome. A variety of batch acquisition functions permit parallelized evaluation of the computational model and support different learning objectives with high efficiency like Bayesian inference, global surrogate development, optimization, etc. Furthermore, libtorch integration supports training, evaluation, and re-training of neural networks and other complex machine learning models in active learning settings. The impacts of these developments are shown on several real applications: (1) nuclear fuel inverse UQ and model inadequacy assessment using the Kennedy O’Hagan framework; (2) uncertainty aware surrogate modeling for additive manufacturing to predict field quantities; (3) nuclear reactor rare events analysis; and (4) complex fluid flow prediction using a global surrogate with quantified prediction uncertainty. Finally, the outlook of MOOSE ProbML is discussed for both outer-loop and inner-loop computations in the broad view to accelerate fuels and materials qualification, address gaps in knowledge and data, and assess new reactor/fuel systems.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Robust Multi-fidelity Bayesian Optimization with Deep Kernel and Partition

Multi-fidelity Bayesian optimization (MFBO) is a powerful approach that utilizes lowfidelity, cost-effective sources to expedite the exploration and exploitation of a high-fidelity objective function. Existing MFBO methods with theoretical foundations either lack justification for performance improvements over single-fidelity optimization or rely on strong assumptions about the relationships between fidelity sources to construct surrogate models and direct queries to low-fidelity sources. To mitigate the dependency on cross-fidelity assumptions while maintaining the advantages of low-fidelity queries, we introduce a random sampling and partition-based MFBO framework with deep kernel learning. This framework is robust to cross-fidelity model misspecification and explicitly illustrates the benefits of low-fidelity queries. Our results demonstrate that the proposed algorithm effectively manages complex cross-fidelity relationships and efficiently optimizes the target fidelity function.

Zhang, Fengxue [University of Chicago, Illinois, U↗

ExtremeMETA: High-speed Lightweight Image Segmentation Model by Remodeling Multi-channel Metamaterial Imagers

Deep neural networks (DNNs) have heavily relied on traditional computational units, such as CPUs and GPUs. However, this conventional approach brings significant computational burden, latency issues, and high power consumption, limiting their effectiveness. This has sparked the need for lightweight networks such as ExtremeC3Net. Meanwhile, there have been notable advancements in optical computational units, particularly with metamaterials, offering the exciting prospect of energy-efficient neural networks operating at the speed of light. Yet, the digital design of metamaterial neural networks (MNNs) faces precision, noise, and bandwidth challenges, limiting their application to intuitive tasks and low-resolution images. In this study, we proposed a large kernel lightweight segmentation model, ExtremeMETA. Based on ExtremeC3Net, our proposed model, ExtremeMETA maximized the ability of the first convolution layer by exploring a larger convolution kernel and multiple processing paths. With the large kernel convolution model, we extended the optic neural network application boundary to the segmentation task. To further lighten the computation burden of the digital processing part, a set of model compression methods was applied to improve model efficiency in the inference stage. The experimental results on three publicly available datasets demonstrated that the optimized efficient design improved segmentation performance from 92.45 to 95.97 on mIoU while reducing computational FLOPs from 461.07 MMacs to 166.03 MMacs. The large kernel lightweight model ExtremeMETA showcased the hybrid design’s ability on complex tasks.

large convolution kernel↗

Transverse momentum dependent PDFs at N3LO

We compute the quark and gluon transverse momentum dependent parton distribution functions at next-to-next-to-next-to-leading order (N 3 LO) in perturbative QCD. Our calculation is based on an expansion of the differential Drell-Yan and gluon fusion Higgs production cross sections about their collinear limit. This method allows us to employ cutting edge multiloop techniques for the computation of cross sections to extract these universal building blocks of the collinear limit of QCD. The corresponding perturbative matching kernels for all channels are expressed in terms of simple harmonic polylogarithms up to weight five. As a byproduct, we confirm a previous computation of the soft function for transverse momentum factorization at N 3 LO. Our results are the last missing ingredient to extend the q T subtraction methods to N 3 LO and to obtain resummed q T spectra at N 3 LL' accuracy both for gluon as well as for quark initiated processes.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

On the Efficient Evaluation of the Exchange Correlation Potential on Graphics Processing Unit Clusters

The predominance of Kohn–Sham density functional theory (KS-DFT) for the theoretical treatment of large experimentally relevant systems in molecular chemistry and materials science relies primarily on the existence of efficient software implementations which are capable of leveraging the latest advances in modern high-performance computing (HPC). With recent trends in HPC leading toward increasing reliance on heterogeneous accelerator-based architectures such as graphics processing units (GPU), existing code bases must embrace these architectural advances to maintain the high levels of performance that have come to be expected for these methods. In this work, we purpose a three-level parallelism scheme for the distributed numerical integration of the exchange-correlation (XC) potential in the Gaussian basis set discretization of the Kohn–Sham equations on large computing clusters consisting of multiple GPUs per compute node. In addition, we purpose and demonstrate the efficacy of the use of batched kernels, including batched level-3 BLAS operations, in achieving high levels of performance on the GPU. We demonstrate the performance and scalability of the implementation of the purposed method in the NWChemEx software package by comparing to the existing scalable CPU XC integration in NWChem.

97 MATHEMATICS AND COMPUTING↗

Multichannel Analysis of Surface Waves Accelerated (MASWAccelerated): Software for efficient surface wave inversion using MPI and GPUs

Multichannel Analysis of Surface Waves (MASW) is a technique frequently used in geotechnical engineering and engineering geophysics to infer 1D layered models of seismic shear wave velocities in the top tens to hundreds of meters of the subsurface. We aim to accelerate MASW calculations by capitalizing on modern computer hardware available in the workstations of most engineers: multiple cores and graphics processing units (GPUs). We propose new parallel and GPU accelerated algorithms for computing 1D MASW inversion, and provide software implementations in C using Message Passing Interface (MPI) and CUDA. These algorithms take advantage of sparsity that arises in the problem, and the work balance between processes considers typical data trends. We compare our methods to an existing open source Matlab MASW tool. Our serial C implementation achieves a 2x speedup over the Matlab software, and we continue to see improvements by parallelizing the problem with MPI. Here we see nearly perfect strong and weak scaling for uniform data, and improve strong scaling for realistic data by repartitioning the problem to process mapping. By utilizing GPUs available on most modern workstations, we observe an additional 1.3x speedup over the serial C implementation on the first use of the method. We typically repeatedly evaluate theoretical dispersion curves as part of an optimization procedure, and on the GPU the kernel can be cached for faster reuse on later runs. We observe a 3.2x speedup on the cached GPU runs compared to the serial C runs. This work is the first open-source parallel or GPU-accelerated software tool for MASW imaging, and should enable geotechnical engineers to fully utilize all computer hardware at their disposal.

58 GEOSCIENCES↗