Engineering Papers⌕ Search

Engineering topics

Mullowney, Paul

Publications and source records attributed to Mullowney, Paul.

The Pele Simulation Suite for Reacting Flows at Exascale

In this work, we present the Pele suite of software tools for compressible and incompressible reacting flows. The Pele suite leverages several different libraries, notably AMReX and SUNDIALS, to achieve performance portability on heterogeneous computing architectures across the supercomputing landscape. The Pele suite is comprised of PeleC, a compressible reacting flow block-structured adaptive mesh refinement solver, PeleLMeX, a low-Mach number reacting flow block-structured adaptive mesh refinement solver, Pele-Physics, a library for transport, thermodynamics, finite rate chemistry, soot, spray and radiation physics. The objective of this paper is (i) to present the code development efforts necessary to achieve highly effective and scalable applications for exascale machines and (ii) to detail the performance results of the Combustion-Pele project applications on Oak Ridge National Laboratory's Frontier. We show good weak and strong scaling results for both PeleC and PeleLMeX up to more than 50 billion cells on more than 4096 Frontier graphics processing unit nodes. We also present a capability demonstration simulation of a dual-fuel pulse compression ignition engine (six adaptive mesh refinement levels, and 60 billion cells or 2.1 trillion degrees of freedom) on Frontier, to date one of the largest simulations performed on the first exascale-class supercomputer.

adaptive mesh refinement↗

ExaWind: Open‐source CFD for hybrid‐RANS/LES geometry‐resolved wind turbine simulations in atmospheric flows

Abstract Predictive high‐fidelity modeling of wind turbines with computational fluid dynamics, wherein turbine geometry is resolved in an atmospheric boundary layer, is important to understanding complex flow accounting for design strategies and operational phenomena such as blade erosion, pitch‐control, stall/vortex‐induced vibrations, and aftermarket add‐ons. The biggest challenge with high‐fidelity modeling is the realization of numerical algorithms that can capture the relevant physics in detail through effective use of high‐performance computing. For modern supercomputers, that means relying on GPUs for acceleration. In this paper, we present ExaWind, a GPU‐enabled open‐source incompressible‐flow hybrid‐computational fluid dynamics framework, comprising the near‐body unstructured grid solver Nalu‐Wind, and the off‐body block‐structured‐grid solver AMR‐Wind, which are coupled using the Topology Independent Overset Grid Assembler. Turbine simulations employ either a pure Reynolds‐averaged Navier–Stokes turbulence model or hybrid turbulence modeling wherein Reynolds‐averaged Navier–Stokes is used for near‐body flow and large eddy simulation is used for off‐body flow. Being two‐way coupled through overset grids, the two solvers enable simulation of flows across a huge range of length scales, for example, 10 orders of magnitude going from O(μm) boundary layers along the blades to O(10 km) across a wind farm. In this paper, we describe the numerical algorithms for geometry‐resolved turbine simulations in atmospheric boundary layers using ExaWind. We present verification studies using canonical flow problems. Validation studies are presented using megawatt‐scale turbines established in literature. Additionally presented are demonstration simulations of a small wind farm under atmospheric inflow with different stability states.

17 WIND ENERGY↗

Scaled ILU Smoothers for Navier-Stokes Pressure Projection

Incomplete LU (ILU) smoothers are effective in the algebraic multigrid (AMG) V-cycle for reducing high-frequency components of the error. However, the requisite direct triangular solves are comparatively slow on GPUs. Previous work has demonstrated the advantages of Jacobi iteration as an alternative to direct solution of these systems. Depending on the threshold and fill-level parameters chosen, the factors can be highly nonnormal and Jacobi is unlikely to converge in a low number of iterations. We demonstrate that row scaling can reduce the departure from normality, allowing us to replace the inherently sequential solve with a rapidly converging Richardson iteration. There are several advantages beyond the lower compute time. Scaling is performed locally for a diagonal block of the global matrix because it is applied directly to the factor. Further, an ILUT Schur complement smoother maintains a constant GMRES iteration count as the number of MPI ranks increases, and thus parallel strong-scaling is improved. Our algorithms have been incorporated into hypre, and we demonstrate improved time to solution for linear systems arising in the Nalu-Wind and PeleLM pressure solvers. For large problem sizes, GMRES+AMG executes at least five times faster when using iterative triangular solves compared with direct solves on massively parallel GPUs.

algebraic multigrid↗

ExaWind at NREL: Upping the Ante

The objective of the ExaWind component of the Exascale Computing Project is to deliver many-turbine blade-resolved simulations in complex terrain. These simulations bring new challenges to both compute and analysis of the resulting data. In this paper/video, we visually explore the impact of ExaWind on wind simulations through two studies of a small wind farm under two atmospheric conditions. We then turn to analysis and review tools that visualization researchers at NREL use to answer the challenges that ExaWind brings.

collaborative visualization↗

Experiences Readying Applications for Exascale

The advent of Exascale computing invites an assessment of existing best practices for developing application readiness on the world's largest supercomputers. This work details observations from the last four years in preparing scientific applications to run on the Oak Ridge Leadership Computing Facility's (OLCF) Frontier system. This paper addresses a range of topics in software including programmability, tuning, and portability considerations that are key to moving applications from existing systems to future installations. A set of representative workloads provides case studies for general system and software testing. We evaluate the use of early access systems for development across several generations of hardware. Finally, we discuss how best practices were identified and disseminated to the community through a wide range of activities including user-guides and trainings. We conclude with recommendations for ensuring application readiness on future leadership computing systems.

exascale↗

Experiences readying applications for Exascale

The advent of Exascale computing invites an assessment of existing best practices for developing application readiness on the world's largest supercomputers. This work details observations from the last four years in preparing scientific applications to run on the Oak Ridge Leadership Computing Facility's (OLCF) Frontier system. This paper addresses a range of topics in software including programmability, tuning, and portability considerations that are key to moving applications from existing systems to future installations. A set of representative workloads provides case studies for general system and software testing. We evaluate the use of early access systems for development across several generations of hardware. Finally, we discuss how best practices were identified and disseminated to the community through a wide range of activities including user-guides and trainings. We conclude with recommendations for ensuring application readiness on future leadership computing systems.

Gottiparthi, Kalyan↗