Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “kernel”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32

Higher-order particle representation for particle-in-cell simulations

In this paper we present an alternative approach to the representation of simulation particles for unstructured electrostatic and electromagnetic PIC simulations. In our modified PIC algorithm we represent particles as having a smooth shape function limited by some specified finite radius, r 0 . A unique feature of our approach is the representation of this shape by surrounding simulation particles with a set of virtual particles with delta shape, with fixed offsets and weights derived from Gaussian quadrature rules and the value of r 0 . As the virtual particles are purely computational, they provide the additional benefit of increasing the arithmetic intensity of traditionally memory bound particle kernels. The modified algorithm is implemented within Sandia National Laboratories' unstructured EMPIRE-PIC code, for electrostatic and electromagnetic simulations, using periodic boundary conditions. We show results for a representative set of benchmark problems, including electron orbit, a transverse electromagnetic wave propagating through a plasma, numerical heating, and a plasma slab expansion. In this work, good error reduction across all of the chosen problems is achieved as the particles are made progressively smoother, with the optimal particle radius appearing to be problem-dependent.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Flow and transport in three-dimensional discrete fracture matrix models using mimetic finite difference on a conforming multi-dimensional mesh

Here, we present a comprehensive workflow to simulate single-phase flow and transport in fractured porous media using the discrete fracture matrix approach. The workflow has three primary parts: (1) a method for conforming mesh generation of and around a three-dimensional fracture network, (2) the discretization of the governing equations using a second-order mimetic finite difference method, and (3) implementation of numerical methods for high-performance computing environments. A method to create a conforming Delaunay tetrahedralization of the volume surrounding the fracture network, where the triangular cells of the fracture mesh are faces in the volume mesh, that addresses pathological cases which commonly arise and degrade mesh quality is also provided. Our open-source subsurface simulator uses a hierarchy of process kernels (one kernel per physical process) that allows for both strong and weak coupling of the fracture and matrix domains. We provide verification tests based on analytic solutions for flow and transport, as well as numerical convergence. We also provide multiple expositions of the method in complex fracture networks. In the first example, we demonstrate that the method is robust by considering two scenarios where the fracture network acts as a barrier to flow, as the primary pathway, or offers the same resistance as the surrounding matrix. In the second test, flow and transport through a three-dimensional stochastically generated network containing 257 fractures is presented.

97 MATHEMATICS AND COMPUTING↗

Sharp front tracking with geometric interface reconstruction

Here, this paper presents a novel sharp front-tracking method designed to address limitations in classical front-tracking approaches, specifically their reliance on smooth interpolation kernels and extended stencils for coupling the front and fluid mesh. In contrast, the proposed method employs exclusively sharp, localized interpolation and spreading kernels, restricting the coupling to the interfacial fluid cells–those containing the interface/front. This localized coupling is achieved by integrating a divergence-preserving velocity interpolation method with a piecewise parabolic interface calculation (PPIC) and a polyhedron intersection algorithm to compute the indicator function and local interface curvature. Surface tension is computed using the Continuum Surface Force (CSF) method, maintaining consistency with the sharp representation. Additionally, we propose an efficient local roughness smoothing implementation to account for surface mesh undulations, which is easily applicable to any triangulated surface mesh. Building on our previous work, the primary innovation of this study lies in the localization of the coupling for both the indicator function and surface tension calculations. By reducing the interface thickness on the fluid mesh to a single cell, as opposed to the 4–5 cell spans typical in classical methods, the proposed sharp front-tracking method achieves a highly localized and accurate representation of the interface. This sharper representation mitigates parasitic currents and improves force balancing, making it particularly suitable for scenarios where the interface plays a critical role, such as microfluidics, fluid-fluid interactions, and fluid-structure interactions. The proposed method is comprehensively validated and tested on canonical interfacial flow problems, including stationary and translating Laplace equilibria, oscillating droplets, and rising bubbles. The presented results demonstrate that the sharp front-tracking method significantly outperforms the classical approach in terms of accuracy, stability, and computational efficiency. Notably, parasitic currents are reduced by approximately two orders of magnitude and stable results are obtained for parameter ranges where classical front tracking fails to converge.

42 ENGINEERING↗

Contributions to the mechanistic understanding of the microstructural evolution in irradiated U-Mo dispersion fuel

Here, advanced microstructural characterization techniques, such as scanning electron microscopy (SEM) and scanning transmission electron microscopy - energy dispersive x-ray spectroscopy (STEM-EDS), were used to interpret the fuel microstructure evolution and fission products behavior in U-Mo dispersion fuel irradiated in the Advanced Test Reactor (ATR) as part of the European Mini-Plate Irradiation Experiment (EMPIrE) test. The larger as-fabricated fuel grain size achieved by heat-treating the U-Mo powder resulted in slower high burnup structure (HBS) development and reduced fission gas porosity. Slower HBS kinetics was observed at the fuel kernels’ periphery, which contained smaller and less fission gas bubbles at all fission densities (FDs) investigated and was attributed to a locally reduced damage density and fission products concentration, as corroborated with Monte Carlo simulations. The non-refined grains at the fuel kernel periphery hosted a perfectly ordered fission Gas Bubble Superlattice (GBS) up to 6.3 × 10 21 fissions/cm 3 . Nano-scale STEM-EDS analysis presented in this study provided useful information on the GBS characteristic morphology and evolution in U-Mo fuel. The concentration of fission gas in the GBS progressively increased with FD, pointing to an evolution of the nanobubble pressure status with irradiation. A possible connection between the GBS collapse and HBS onset is proposed for which there exists a threshold in the misorientation of the refined sub-grains above which the GBS stability during irradiation is no longer preserved, resulting in the GBS collapse.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Optimizing the hypre solver for manycore and GPU architectures

The solution of large-scale combustion problems with codes such as Uintah on modern computer architectures requires the use of multithreading and GPUs to achieve performance. Uintah uses a low-Mach number approximation that requires iteratively solving a large system of linear equations. The Hypre iterative solver has solved such systems in a scalable way for Uintah, but the use of OpenMP with Hypre leads to at least slowdown due to OpenMP overheads. The proposed solution uses the MPI Endpoints within Hypre, where each team of threads acts as a different MPI rank. This approach minimizes OpenMP synchronization overhead and performs as fast or (up to 1.44) faster than Hypre's MPI-only version, and allows the rest of Uintah to be optimized using OpenMP. The profiling of the GPU version of Hypre shows the bottleneck to be the launch overhead of thousands of micro-kernels. The GPU performance was improved by fusing these micro-kernels and was further optimized by using Cuda-aware MPI, resulting in an overall speedup of 1.16—1.44 compared to the baseline GPU implementation. The above optimization strategies were published in the International Conference on Computational Science 2020 [1]. This work extends the previously published research by carrying out the second phase of communication-centered optimizations in Hypre to improve its scalability on large-scale supercomputers. Additionally, this includes an efficient non-blocking inter-thread communication scheme, communication-reducing patch assignment, and expression of logical communication parallelism to a new version of the MPICH library that utilizes the underlying network parallelism [2]. The above optimizations avoid communication bottlenecks previously observed during strong scaling and improve performance by up to 2 on 256 nodes of Intel Knight's Landing processor.

97 MATHEMATICS AND COMPUTING↗

A Fine-grained Asynchronous Bulk Synchronous parallelism model for PGAS applications

The Partitioned Global Address Space (PGAS) model is well suited for executing irregular applications on cluster-based systems, due to its efficient support for short, one-sided messages. Separately, the actor model has been gaining popularity as a productive asynchronous message-passing approach for distributed objects in enterprise and cloud computing platforms, typically implemented in languages such as Erlang, Scala or Rust. To the best of our knowledge, there has been no past work on using the actor model to deliver both productivity and scalability to irregular PGAS applications with large number of small messages. In this paper, we introduce a new programming system for PGAS applications, in which point-to-point remote operations can be expressed as fine-grained asynchronous actor messages. In our approach, the programmer does not need to worry about programming complexities related to message aggregation and termination detection. Our approach can be viewed as extending the classical Bulk Synchronous Parallelism model with fine-grained asynchronous communications within a phase or superstep. Here, we believe that our approach offers a desirable point in the productivity-performance space for PGAS applications, with more scalable performance and higher productivity relative to past approaches. Specifically, for seven irregular mini-applications from the Bale Kernels and three graph kernels executed using 2048 cores in the NERSC Cori system, our approach shows geometric mean performance improvements of ≥ 20X relative to standard PGAS versions (UPC and OpenSHMEM) while maintaining comparable productivity to those versions.

97 MATHEMATICS AND COMPUTING↗

Fast GPU 3D diffeomorphic image registration

3D image registration is one of the most fundamental and computationally expensive operations in medical image analysis. Here, we present a mixed-precision, Gauss–Newton–Krylov solver for diffeomorphic registration of two images. Our work extends the publicly available CLAIRE library to GPU architectures. Despite the importance of image registration, only a few implementations of large deformation diffeomorphic registration packages support GPUs. Our contributions are new algorithms to significantly reduce the run time of the two main computational kernels in CLAIRE: calculation of derivatives and scattered-data interpolation. Additionally, we deploy (i) highly-optimized, mixed-precision GPU-kernels for the evaluation of scattered-data interpolation, (ii) replace Fast-Fourier-Transform (FFT)-based first-order derivatives with optimized 8th-order finite differences, and (iii) compare with state-of-the-art CPU and GPU implementations. As a highlight, we demonstrate that we can register clinical images in less than 6 s on a single NVIDIA Tesla V100. This amounts to over 20 speed-up over the current version of CLAIRE and over 30 speed-up over existing GPU implementations.

97 MATHEMATICS AND COMPUTING↗

Machine learning for the redox potential prediction of molecules in organic redox flow battery

Here, organic redox flow batteries (ORFB) are recognized as an innovative technology for the large-scale storage of renewable energy. The redox potential of organic redox-active molecules plays a vital role in their performance. Advanced screening techniques like high-throughput experiment and machine learning (ML) have significantly enhanced organic material performance and transformed the field of ORFB. However, the scarcity of experimental data poses a considerable challenge for ML model development in this domain. In our study, we developed lightweight graph-based Gaussian process regression (GPR) models with GPU-accelerated marginalized graph kernel and hybrid kernel to predict the redox potentials of organic redox-active molecules for ORFBs, specifically focusing on small datasets. To evaluate model accuracy, we created a new experimental database of organic redox-active molecules by the data from hundreds of published papers and assembled previous computational datasets. We also considered some key parameters, such as pH conditions and solvent type, to assess their impact on redox potential prediction. Our GPR model predicted redox potentials with high accuracy across all datasets using minimal training data. The study provides powerful tools for molecule screening and design and delivers valuable guidance on designing training datasets for costly experiments.

25 ENERGY STORAGE↗

Evaluation of the spatial self-shielding impact for TRISO-based nuclear fuel depletion

Reactor physics analyses of nuclear cores with nuclear fuel concepts containing tristructural isotropic (TRISO) particles, such as pebbles or compact fuel elements, rely on various degrees of simplification to keep these highly heterogeneous problems computationally tractable. One such limitation regards the level of spatial discretization employed during burnup calculations, where traditionally only a limited number of spatial zones are modeled at the full core level and assume that the spectrum is constant within the fuel elements and TRISO particles in this depletion zone. This type of assumption neglects the impact of spatial self-shielding effect within the kernels (microscale level) as well as within the compact or pebbles (mesoscale level). Furthermore, the Monte Carlo code Serpent 2 contains many relevant features for efficiently modeling this type of geometry, including a collision-based domain decomposition intended for very large burnup calculations, which we leveraged for this work to quantify the impact of capturing neutron flux variations occurring at the micro- and mesoscale level on a series of high-temperature gas-cooled reactor fuel element depletion problems. While spatial self-shielding is observed at both scales, with differences from a volume-averaged burnup of ±7% within the kernels and ±2% between TRISO particles within the fuel element, the conjugated effect on nuclide inventories and multiplication factor are negligible, hence confirming that assuming a single average spectrum value may be sufficient for most applications.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Achieving performance portability in Gaussian basis set density functional theory on accelerator based architectures in NWChemEx

The numerical integration of the exchange–correlation (XC) potential is one of the primary computational bottlenecks in Gaussian basis set Kohn–Sham density functional theory (KS-DFT). To achieve optimal performance and accuracy, care must be taken in this numerical integration to preserve local sparsity as to allow for near linear weak scaling with system size. This leads to an integration scheme with several performance critical kernels which must be hand optimized for each architecture of interest. As the set of available accelerator hardware goes more diverse, a key challenge for developers of KS-DFT software is to maintain performance portability across a wide range of computational architectures. In this article, we examine a modular software design pattern which decouples the implementation details of performance critical kernels from the expression of high-level algorithmic workflows in a device-agnostic language such as C++; thus allowing for developers to target existing and emerging accelerator hardware within a single code base. We consider the efficacy of such a design pattern in the numerical integration of the XC potential by demonstrating its ability to achieve performance portability across a set of accelerator architectures which are representative of those on current and future U.S. Department of Energy Leadership Computing Facilities.

97 MATHEMATICS AND COMPUTING↗

High performance sparse multifrontal solvers on modern GPUs

Here, we have ported the numerical factorization and triangular solve phases of the sparse direct solver STRUMPACK to GPU. STRUMPACK implements sparse LU factorization using the multifrontal algorithm, which performs most of its operations in dense linear algebra operations on so-called frontal matrices of various sizes. Our GPU implementation off-loads these dense linear algebra operations, as well as the sparse scatter–gather operations between frontal matrices. For the larger frontal matrices, our GPU implementation relies on vendor libraries such as cuBLAS and cuSOLVER for NVIDIA GPUs and rocBLAS and rocSOLVER for AMD GPUs. For the smaller frontal matrices we developed custom CUDA and HIP kernels to reduce kernel launch overhead. Overall, high performance is achieved by identifying submatrix factorizations corresponding to sub-trees of the multifrontal assembly tree which fit entirely in GPU memory. The multi-GPU setting uses SLATE (Software for Linear Algebra Targeting Exascale) as a modern GPU-aware replacement for ScaLAPACK. On 4 nodes of SUMMIT the code runs ~10X faster when using all 24 V100 GPUs compared to when it only uses the 168 POWER9 cores. On 8 SUMMIT nodes, using 48 V100 GPUs, the sparse solver reaches over 50TFlop/s. Compared to SuperLU, on a single V100, for a set of 17 matrices our implementation is faster for all but one matrix, and is on average 5X (median 4X) faster

97 MATHEMATICS AND COMPUTING↗

High volume packing fraction TRISO-based fuel in light water reactors

We report that for a decade, fully ceramic microencapsulated (FCM) fuel, containing tri-structural isotropic (TRISO) fuel particles in a silicon carbide (SiC) matrix, has been investigated as an accident-tolerant fuel for light water reactors (LWRs). Other examples exist of TRISO-based concepts for LWR fuels with different matrix materials. Previous studies assumed TRISO particle volume packing of approximately 0.44 in SiC (or another) matrix, the highest realistic packing fractions possible with conventional manufacturing. Recent advances in advanced manufacturing have yielded the development and demonstration of a fuel form that consists of conventionally manufactured TRISO particles in a 3D-printed SiC matrix with significantly higher possible TRISO packing fractions (0.5–0.7). This increased uranium loading enhances the viability of using TRISO-based particle fuel forms in LWRs. The viability of high-packing-fraction TRISO-based particle fuel forms in LWRs is assessed from the perspective of fuel cycle length, achievable fuel burnup, reactivity coefficients, and fuel cycle performance. Higher-packing-fraction TRISO-based fuel enables either longer cycle lengths (by ~25% at a packing fraction of 0.55 relative to 0.44) at a constant enrichment or decreased enrichments (by ~25% at a packing fraction of 0.55 relative to 0.44) at a constant cycle length. Studies of different fuel kernel types (uranium nitride, uranium oxycarbide, and uranium carbide) yield similar results, although the cycle length of uranium oxycarbide is shorter than for uranium nitride or uranium carbide (due to the lower density of uranium oxide). This work also characterized the production of 14 C resulting from neutron absorption in 14 N during operation for uranium mononitride fuel kernels; the ratio of 14 C/N was 1–2 at. % at discharge. For the fuel cycle evaluation, the activity of spent nuclear fuel and high-level waste at 100 and 100,000 years was lower for high-packing-fraction fuels than for conventional LWR fuel. Environmental impact metrics were similar overall, but higher on the front end of the fuel cycle and lower on the back end of the fuel cycle. Reactivity coefficients of higher-packing-fraction TRISO-based fuel were reasonable compared with those of conventional fuels.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Data-Efficient Strategies for Probabilistic Voltage Envelopes under Network Contingencies

This work presents an efficient data-driven method to construct probabilistic voltage envelopes (PVE) using power flow learning in grids with network contingencies. First, a network-aware Gaussian process (GP) termed Vertex-Degree Kernel (VDK-GP), developed in prior work, is used to estimate voltage–power functions for a few network configurations. The paper introduces a novel multi-task vertex degree kernel (MT-VDK) that amalgamates the learned VDK-GPs to determine power flows for unseen networks, with a significant reduction in the computational complexity and hyperparameter requirements compared to alternate approaches. Simulations on the IEEE 30-Bus network demonstrate the retention and transfer of power flow knowledge in both N-1 and N-2 contingency scenarios. The MT-VDK-GP approach achieves over 50 % reduction in mean prediction error for novel N-1 contingency network configurations in low training data regimes (50–250 samples) over VDK-GP. Additionally, MT-VDK-GP outperforms a hyper-parameter based transfer learning approach in over 75 % of N-2 contingency network structures, even without historical N-2 outage data. Furthermore, the proposed method demonstrates the ability to achieve PVEs using sixteen times fewer power flow solutions compared to Monte-Carlo sampling-based methods.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Seeing the whole picture: Methods for getting the most from micro X-ray computed tomography of TRISO nuclear fuel particles

Tristructural isotropic (TRISO) coated fuel particles are a nuclear fuel form under extensive study for use in advanced nuclear reactor concepts. TRISO fuels are subjected to high temperature neutron irradiations and then examined to assess their performance by determining fission product retention and studying morphological changes. Micro X-ray computed tomography is one method of nondestructively studying the effects of TRISO performance. This work addresses the need for image processing to remove X-ray tomographic reconstruction artifacts that prevent the study of TRISO features, as the TRISO particles’ high Z kernel can introduce metal artifacts that degrade the image quality in the surrounding low Z coating layers. These metal artifacts were reduced by imaging the TRISO particles with both high- and low-energy X-rays and applying a mask to the radiographs obtained with low-energy X-rays to digitally remove the dense fuel kernel region. These masked radiographs were then used to produce a tomographic reconstruction which was combined with the tomographic reconstruction of the high-energy data. This enabled the relatively-low-density TRISO buffer layer to be examined in more detail, providing information on irradiation induced dimensional changes of the coatings. This methodology, which helps see the full picture of a TRISO particle, is not limited to nuclear fuels but can be applied to systems that contain highly attenuating material surrounded by less dense materials.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Non-locality of mean scalar transport in two-dimensional Rayleigh–Taylor instability using the macroscopic forcing method

The importance of non-locality of mean scalar transport in two-dimensional Rayleigh–Taylor Instability (RTI) is investigated. The macroscopic forcing method is utilized to measure spatio-temporal moments of the eddy diffusivity kernel representing passive scalar transport in the ensemble averaged fields. Presented in this work are several studies assessing the importance of the higher-order moments of the eddy diffusivity, which contain information about non-locality, in models for RTI. First, it is demonstrated through a comparison of leading-order models that a purely local eddy diffusivity is insufficient to capture the mean field evolution of the mass fraction in RTI. Therefore, higher-order moments of the eddy diffusivity operator are not negligible. Models are then constructed by utilizing the measured higher-order moments. It is demonstrated that an explicit operator based on the Kramers–Moyal expansion of the eddy diffusivity kernel is insufficient. An implicit operator construction that matches the measured moments is shown to offer improvements relative to the local model in a converging fashion.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Spatial Signatures of Electron Correlation in Least-Squares Tensor Hypercontraction

Least Squares Tensor Hypercontraction (LS-THC) has received some attention in recent years as an approach to reduce the significant computational costs of wavefunc- tion based methods in quantum chemistry. However, previous work has demonstrated that the LS-THC factorization performs disproportionately worse in the description of wavefunction components (e.g. cluster amplitudes T 2 ) than Hamiltonian compo- nents (e.g. electron repulsion integrals (pq|rs)). This work develops novel theoretical methods to study the source of these errors in the context of the real-space T 2 kernel, and reports, for the first time, the existence of a “correlation feature” in the errors of the LS-THC representation of the “exchange-like” correlation energy EX and T 2 that is remarkably consistent across ten molecular species, three correlated wavefunctions, and four basis sets. This correlation feature portends the existence of a “pair-point kernel” missing in the usual LS-THC representation of the wavefunction, which critically depends upon pairs of grid points situated close to atoms and with inter-pair distances between one and two Bohr radii. These findings point the way for future LS-THC developments to address these shortcomings.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Source of Bright Near-Infrared Luminescence in Gold Nanoclusters

Gold nanoclusters with near-infrared (NIR) photoluminescence (PL) have great potential as sensing and imaging materials in biomedical and bioimaging applications. In this work, Au 21 (S-Adm) 15 and Au 38 S 2 (S-Adm) 20 are used to unravel the underlying mechanisms for the improved quantum yields (QY), large Stokes shifts and long PL lifetimes in gold nanoclusters. Both nanoclusters show decent PL QY. In particular, the Au 38 S 2 (S-Adm) 20 nanocluster shows a bright NIR PL at 900 nm with QY up to 15% in normal solvents (such as toluene) at ambient conditions. The relatively lower QY for Au 21 (S-Adm) 15 (4%) compared to Au 38 S 2 (S-Adm) 20 is attributed to the lowest-lying excited state being symmetry-disallowed, as evidenced by the pressure-dependent anti-spectral shift of the absorption spectra compared to PL. Yet, Au 21 (S-Adm) 15 maintains some emissive properties due to a nearby symmetry-allowed excited state. Furthermore, our results show that suppression of non-radiative decay due to the surface “lock rings” which encircle the Au kernel and the surface “lock atoms” which bridge the fundamental Au-kernel units (e.g., tetrahedra, icosahedra, etc.) is the key to obtain high QYs in gold nanoclusters. Here, the complicated excited-state processes and the small absorption coefficient of the band-edge transition lead to the large Stokes shifts and the long PL lifetimes that are widely observed in gold nanoclusters.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Assessing Global and Local Radiative Feedbacks Based on AGCM Simulations for 1980–2014/2017

In order to avoid contamination of unobservable and uncertain effective radiative forcing (ERF) on the diagnosis of radiative feedbacks based on short-term climate variability, we verify the Kernel-Gregory feedback calculation method using atmospheric model experiments with prescribed ERF. We show that both clear-sky radiative fluxes and all-sky radiative feedbacks have a closure between model simulations and kernel derivations. A near-zero global mean net cloud feedback is found in the simulation with prescribed ERF, which results in a more negative global net climate feedback, -2 W m -2 K -1 . Consistent with AMIP6 ensemble mean results, the lapse rate feedback is the largest contributor among all feedbacks to temperature amplification over the three poles (Arctic, Antarctic and Tibetan Plateau), followed by surface albedo feedback and Planck feedback deviation from its global mean. Except for the higher surface albedo feedback in the Antarctic, other feedbacks are almost same between the Arctic and Antarctic.

54 ENVIRONMENTAL SCIENCES↗