A Robust Parallel Adaptive Mesh Refinement Software Library for Unstructured Meshes
The design and development of a software tool for performing parallel adaptive mesh refinement in unstructured computations on multiprocessor systems are described.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
The design and development of a software tool for performing parallel adaptive mesh refinement in unstructured computations on multiprocessor systems are described.
Thermochemical nonequilibrium flow simulation capabilities have been previously implemented, verified, and validated for central processing unit (CPU) systems in NASA’s unstructured-grid computational fluid dynamics solver FUN3D. Many exascale-class high-performance computing systems will rely on graphics processing unit (GPU) architectures for high throughput and energy efficiency; thus, CPU-based scientific computing software unable to effectively utilize these systems must be updated. In this work, we present a CUDA C++ implementation of FUN3D’s thermochemical nonequilibrium flow simulation capabilities targeting NVIDIA Tesla GPUs. An overview of the porting and optimization strategy is described and performance comparisons with other recent architectures are presented. Scaling to thousands of GPUs is demonstrated, yielding computational performance equivalent to that of several million CPU cores. The implementation enables efficient, high-fidelity, scale-resolving simulations of thermochemical nonequilibrium flows for many applications including atmospheric entry, hypersonics, and combustion.
Aerothermodynamic prediction is important for the design and analysis of many aerospace vehicles. Computational Fluid Dynamics (CFD) tools used for these predictions commonly rely on structured shock-aligned grids due to the strong shocks exhibited at high speeds. This work details the incorporation of an improved HLLE++ scheme for both perfect gas and thermochemical nonequilibrium flows which robustly captures shock waves on both non-shock-aligned and shock-aligned grids. The HLLE++ scheme reduces the impact of eigenvalue limiting on flow solutions which is required for Roe's scheme for high Mach number flows. An adapted grid approach is also demonstrated using mixed-element unstructured grids to improve heat prediction. Thin prismatic boundary layers are used with tetrahedra in the farfield and for capturing shock waves. Various results are presented for hypersonic flows over blunt bodies including cylinders, spheres, and capsules.
Aerothermodynamic prediction is important for the design and analysis of many aerospace vehicles. Computational Fluid Dynamics (CFD) tools used for these predictions commonly rely on structured shock-aligned grids due to the strong shocks exhibited at high speeds. This work details the incorporation of an improved HLLE++ scheme for both perfect gas and thermochemical nonequilibrium flows which robustly captures shock waves on both non-shock-aligned and shock-aligned grids. The HLLE++ scheme reduces the impact of eigenvalue limiting on flow solutions which is required for Roe's scheme for high Mach number flows. An adapted grid approach is also demonstrated using mixed-element unstructured grids to improve heat prediction. Thin prismatic boundary layers are used with tetrahedra in the farfield and for capturing shock waves. Various results are presented for hypersonic flows over blunt bodies including cylinders, spheres, and capsules.
Computational performance of the FUN3D unstructured-grid computational fluid dynamics (CFD) application on massively parallel GPU environments is memory-bound and highly dependent upon efficient reads from and atomic updates to the irregular cell-, edge-, and node-based data structures. In this talk, we present recent efforts into optimizing select performance-critical kernels on NVIDIA Tesla V100 and A100 GPUs and AMD CDNA MI100 GPUs. A novel use of L2 cache residency controls and asynchronous loads into on-chip shared memory are explored on the A100 GPU for the sparse iterative solver, which is dominated by mixed-precision, sparse matrix vector multiplication. Demonstrations show that these methods improve global memory bandwidth utilization by 13.5% on the A100 GPU. Several techniques are also presented that use registers and/or shared memory to facilitate array transposition and aggregation which combine to reduce the frequency and increase the cache efficiency of floating-point atomic updates to the irregular data structures. These methods are demonstrated to improve the kernel throughput by nearly 500% on select kernels on the AMD MI100 over atomic updates directly to global memory. Overall, both V100 and A100 GPUs outperformed the MI100 GPU on kernels dominated by double-precision atomic updates; however, the techniques demonstrated here reduced the performance gap and improved the MI100 performance.
ABSTRACT This paper presents the multi‐species global impurity transport capability developed in a GPU‐accelerated fully 3D unstructured mesh‐based code, GITRm, to simultaneously track multiple impurity species and handle interactions of these impurities with mixed‐material surfaces. Different computational approaches to model particle‐surface interaction or surface response have been developed and compared. Sheath electric field is taken into account by employing a fast distance‐to‐boundary calculation, which is carried out in parallel on distributed or partitioned meshes on multiple GPUs without the need for any inter‐process communication during the simulation. Several example cases, including two for the DIII‐D tokamak, that is, one with the SAS‐V divertor and the other with the collector probes, are used to demonstrate the utility of the current multi‐species capability. For the DIII‐D probe case, the capability of GITRm to resolve the spatial distribution of particles in localized regions, such as diagnostic probes, within non‐axisymmetric tokamak geometries is demonstrated. These simulations involve up to 320 million particles and utilize up to 48 GPUs.
Here, this paper presents a variational multiscale (VMS) based finite element method where the stabilization parameter is computed dynamically. The current dynamic procedure takes in a general structure/form of the stabilization parameter with unknown coefficients and computes them dynamically in a local fashion resulting in a dynamic VMS-based finite element method. Thus, a static stabilization parameter with pre-defined coefficients is not needed. A variational Germano identity (VGI) based local procedure suitable for unstructured meshes is developed to perform the dynamic computation in a local fashion. The local VGI based procedure is applied for each interior vertex in the mesh and unknown coefficients are first determined locally at each vertex, and subsequently, for each element a maximum value is taken over the vertices of the element. To make the current procedure practical, a coarser secondary solution is constructed from the primary coarse-scale solution, which is done locally over a patch of elements around each interior vertex. Further, averaging steps are employed to make the local dynamic procedure robust. Currently, the new dynamic VMS formulation is applied to steady problems governed by the advection-diffusion and incompressible Navier-Stokes equations in both 1D and 2D to demonstrate its efficacy and effectiveness.
Here, we discuss the implementation of a finite element method, used to numerically solve the Euler equations of compressible flows, using an asynchronous runtime system (RTS). The algorithm is implemented for distributed-memory machines, using stationary unstructured 3D meshes, combining data-, and task-parallelism on top of the Charm++ RTS. Charm++’s execution model is asynchronous by default, allowing arbitrary overlap of computation and communication. Task-parallelism allows scheduling parts of an algorithm independently of, or dependent on, each other. Built-in automatic load balancing enables continuous redistribution of computational load by migration of work units based on real-time CPU load measurement. The RTS also features automatic checkpointing, fault tolerance, resilience against hardware failure, and supports power-, and energy-aware computation. We demonstrate scalability up to 25 x 10 9 cells at $\mathscr{O}$10 4 compute cores and the benefits of automatic load balancing for irregular workloads. The full source code with documentation is available at https://quinoacomputing.org.
Geometry deformation due to thermal expansion influences neutron transport in many systems. Studying this phenomenon involves coupling models for neutronics, thermal hydraulics, and solid mechanics. To enable high fidelity modeling of these coupled physics, new capabilities were introduced in Cardinal, coupling OpenMC Monte Carlo particle transport models with MOOSE thermomechanical physics on unstructured moving-mesh geometries. In this work, we present a fully open-source capability leveraging on-the-fly mesh skinning to automatically regenerate OpenMC geometry, which allows multiphysics feedback from temperature, density, and geometry changes. The new capability is verified using an analytic benchmark slab problem, which couples S 2 neutron transport with thermal conduction, convective boundary conditions, Doppler-broadened cross sections, and nonlinear thermal expansion effects along the heated slab. Cardinal reproduces the analytic solutions for the neutron flux, heating, k eff , and temperature with demonstrated convergence in various error terms including mesh resolution and cross section temperature library spacing. For the nominal benchmark conditions and with a fine mesh, maximum relative errors for neutron flux, temperature, and heating are lower than 1%, while errors in integral quantities such as k eff and slab length are within 1 pcm and 48 µm, respectively. This work (i) presents a new numerical approach to thermomechanics coupling with OpenMC models, (ii) is the first (to our knowledge) to utilize a mechanical partial differential equation (PDE) solution to solve the (Griesheimer and Kooreman, 2022) analytic benchmark, and (iii) develops this verified capability within an open-source package.
Abstract The mechanisms and geographic distribution of global tidal dissipation in barotropic tidal models are examined using a high resolution unstructured mesh finite element model. Mesh resolution varies between 2 and 25 km and is especially focused on inner shelves and steep bathymetric gradients. Tidal response sensitivities to bathymetric changes are examined to put into context response sensitivities to frictional processes. We confirm that the Ronne Ice Shelf dramatically affects Atlantic tides but also find that bathymetry in the Hudson Bay system is a critical control. We follow a sequential frictional parameter optimization process and use TPXO9 data‐assimilated tidal elevations as a reference solution. From simulated velocities and depths, dissipation within the global model is estimated and allows us to pinpoint dissipation at high resolution. Boundary layer dissipation is extremely focused with 1.4% of the ocean accounting for 90% of the total. Internal tide friction is much more distributed with 16.7% of the ocean accounting for 90% of the total. Often highly regional dissipation can impact basin‐scale and even ocean wide tides. Optimized boundary layer friction parameters correlate very well with the physical characteristics of the locality with high friction factors associated with energetic tidal regions, deep ocean island chains, and ice covered areas. Global complex M 2 tide errors are 1.94 cm in deep waters. Total global boundary layer and internal tide dissipation are estimated, respectively, at 1.83 and 1.49 TW. This continues the trend in the literature toward attributing more dissipation to internal tides.
Abstract This work presents a new reactive transport framework that combines a powerful geochemistry engine with advanced numerical methods for flow and transport in subsurface fractured porous media. Specifically, the PhreeqcRM interface (developed by the USGS) is used to take advantage of a large library of equilibrium and kinetic aqueous and fluid-rock reactions, which has been validated by numerous experiments and benchmark studies. Fluid flow is modeled by the Mixed Hybrid Finite Element (FE) method, which provides smooth velocity fields even in highly heterogenous formations with discrete fractures. A multilinear Discontinuous Galerkin FE method is used to solve the multicomponent transport problem. This method is locally mass conserving and its second order convergence significantly reduces numerical dispersion. In terms of thermodynamics, the aqueous phase is considered as a compressible fluid and its properties are derived from a Cubic Plus Association (CPA) equation of state. The new simulator is validated against several benchmark problems (involving, e.g., Fickian and Nernst-Planck diffusion, isotope fractionation, advection-dispersion transport, and rock-fluid reactions) before demonstrating the expanded capabilities offered by the underlying FE foundation, such as high computational efficiency, parallelizability, low numerical dispersion, unstructured 3D gridding, and discrete fraction modeling.
We present a general framework for compressing unstructured scientific data with known local connectivity. A common application is simulation data defined on arbitrary finite element meshes. The framework employs a greedy topology preserving reordering of original nodes which allows for seamless integration into existing data processing pipelines. This reordering process depends solely on mesh connectivity and can be performed offline for optimal efficiency. However, the algorithm’s greedy nature also supports on-the-fly implementation. The proposed method is compatible with any compression algorithm that leverages spatial correlations within the data. The effectiveness of this approach is demonstrated on a large-scale real dataset using several compression methods, including MGARD, SZ, and ZFP.
A novel method for pairing surface irradiation and volumetric absorption from Monte Carlo ray tracing to computational heat transfer models is presented. The method is well-suited to directionally and spatially complex concentrated radiative inputs (e.g., solar receivers and reactors). The method employs a generalized algorithm for directly mapping absorbed rays from a Monte Carlo ray tracing model to boundary or volumetric source terms in the computational mesh. The algorithm is compatible with unstructured, two and three-dimensional meshes with varying element shapes. Four case studies were performed on a directly irradiated, windowed solar thermochemical reactor model to validate the method. The method was shown to conserve energy and preserve spatial variation when mapping rays from a Monte Carlo ray tracing model to a computational heat transfer model in ansys fluent.
The SARS-CoV-2 nucleocapsid (N) protein is highly immunogenic, and anti-N antibodies are commonly used as markers for prior infection. While several studies have examined or predicted the antigenic regions of N, these have lacked consensus and structural context. Using COVID-19 patient sera to probe an overlapping peptide array, we identified six public and four private epitope regions across N, some of which are unique to this study. We further report the first deposited X-ray structure of the stable dimerization domain at 2.05 Å as similar to all other reported structures. Structural mapping revealed that most epitopes are derived from surface-exposed loops on the stable domains or from the unstructured linker regions. An antibody response to an epitope in the stable RNA binding domain was found more frequently in sera from patients requiring intensive care. Since emerging amino acid variations in N map to immunogenic peptides, N protein variation could impact detection of seroconversion for variants of concern.
Ume is an open-source collection of data structures for unstructured computational meshes and some simple algorithms that operate on them. These algorithms mimic the memory access patterns of a common class of operations found in several of the computational physics simulation codes developed at Los Alamos National Laboratory. The intent is that Ume can be used by hardware vendors to understand the memory traffic created by complex codes in a simplified environment, and to explore new means of optimization for that traffic. Ume is provided as a source-code C++ library and includes several applications that demonstrate the use of that library.
ASAUM is a C++ library for representing distributed, multi-block structed and unstructured meshes. The library is meant to provide users with the ability to read/write, partition, and query the large distributed meshes that are commonly used in scientific computing.
SAND2021-15058 O This code accompanies the International Conference on Learning Representations (ICLR) paper entitled "Parameterized Pseudo-Differential Operators for Applying CNNs to Unstructured Mesh Data." Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.
The Oak Ridge National Laboratory (ORNL) Second Target Station (STS) neutron production facility is an accelerator driven pulsed neutron source that is currently being actively developed at ORNL. The neutrons are produced by proton-induced spallation reactions. A proton beam of 700 kW power is delivered to a spallation target in short, less than 1 µs long pulses, with 15 Hz repetition rate. The spallation target of ORNL STS is a rotating water-cooled tungsten target with tantalum cladding housed in a stainless-steel shroud. It is divided into 21 segments. These segments become highly activated due to spallation reactions or nuclei transmutation by the emitted neutrons. The radioactive nuclides continue to decay after ceasing operation. The decay dose rates generated from the target segments once they are removed from their operational location within the core vessel must be accurately quantified to determine the shielding configurations of remote handling tools and transport casks and to aid in planning maintenance events. To determine the shielding configurations needed for an activated target segment after ceasing operation, both the hybrid unstructured mesh (UM)/constructive solid geometry (CSG) approach that was previously utilized for STS analyses [1] and the ADVANTG code [2] were used. Even though the ADVANTG code does not include UM capability, the utilization of its advanced variance reduction technique was crucial to accelerate the extremely difficult final photon transport calculation in this analysis. This paper also describes the procedures taken to mitigate the convergence issues that occur when ADVANTG uses a source definition that does not match the source of the final Monte Carlo (MC) calculation. These convergence issues often occur because of inconsistencies between source and transport biasing parameters.