Engineering Papers⌕ Search

Engineering topics

Spies, Lukas

Publications and source records attributed to Spies, Lukas.

Tausch: A halo exchange library for large heterogeneous computing systems using MPI, OpenCL, and CUDA

Exchanging halo data is a common task in modern scientific computing applications and efficient handling of this operation is critical for the performance of the overall simulation. Tausch is a novel header-only library that provides a simple API for efficiently handling these types of data movements. Tausch supports both simple CPU-only systems, but also more complex heterogeneous systems with both CPUs and GPUs. It currently supports both OpenCL and CUDA for communicating with GPGPU devices, and allows for communication between GPGPUs and CPUs. The API allows for drop-in replacement in existing codes and can be used for the communication layer in new codes. This paper provides an overview of the approach taken in Tausch, and a performance analysis that demonstrates expected and achieved performance. Here, we highlight the ease of use and performance with three applications: First Tausch is compared to the halo exchange framework from two Mantevo applications, HPCCG and miniFE, and then it is used to replace a legacy halo exchange library in the flexible multigrid solver framework Cedar.

97 MATHEMATICS AND COMPUTING↗

Comparison of Some RANS Solvers

We will take a look at solving the Reynolds-averaged Navier-Stokes (RANS) equations that are encountered in the context of wind farm performance simulations and optimizations. We will compare some of the more popular ways to solve these equations with a focus on using iterative solvers for the linear solve. We will compare their performance and reliability to a direct solve as we scale the problem both by adding more and parallel resources and by increasing the size of the domain, both in two and three dimensions. There are many strategies that can be applied to solving the RANS equations, some are very efficient, while others are very insensitive to, for example, the Reynolds number. The first contender we will consider is the Pressure-Convection-Diffusion (PCD) preconditioner. Early results suggest that PCD is indeed a very efficient solver, in particular in two dimensions, as long as the Reynold's number remains small. Next we will try to reorder our degrees of freedom such that we can use GMRES with ILU for our linear solve. Another popular choice we will consider for solving the RANS equations is SIMPLE (and its derivatives). For all of our implementations we make use of either FEniCS or Firedrake, basing our work on both existing implementations of some of these solvers while also writing new extensions for others.

CFD↗

Efficient exascale discretizations: High-order finite element methods

Efficient exploitation of exascale architectures requires rethinking of the numerical algorithms used in many large-scale applications. These architectures favor algorithms that expose ultra fine-grain parallelism and maximize the ratio of floating point operations to energy intensive data movement. One of the few viable approaches to achieve high efficiency in the area of PDE discretizations on unstructured grids is to use matrix-free/partially assembled high-order finite element methods, since these methods can increase the accuracy and/or lower the computational time due to reduced data motion. In this paper we provide an overview of the research and development activities in the Center for Efficient Exascale Discretizations (CEED), a co-design center in the Exascale Computing Project that is focused on the development of next-generation discretization software and algorithms to enable a wide range of finite element applications to run efficiently on future hardware. CEED is a research partnership involving more than 30 computational scientists from two US national labs and five universities, including members of the Nek5000, MFEM, MAGMA and PETSc projects. We discuss the CEED co-design activities based on targeted benchmarks, miniapps and discretization libraries and our work on performance optimizations for large-scale GPU architectures. We also provide a broad overview of research and development activities in areas such as unstructured adaptive mesh refinement algorithms, matrix-free linear solvers, high-order data visualization, and list examples of collaborations with several ECP and external applications.

97 MATHEMATICS AND COMPUTING↗