Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hypre”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

hypredrive: high-level interface for solving linear systems with hypre

This software introduces a high-level interface designed to simplify solving linear systems using hypre, a renowned library for such computational challenges. It is crafted to be accessible and user-friendly, making the powerful capabilities of hypre available to a broader audience without requiring in-depth technical knowledge. The interface is characterized by its use of YAML for input, a format celebrated for its structured yet straightforward readability. This choice ensures that users can easily configure the software to meet their specific needs. Additionally, the software boasts an intuitive API that encapsulates hypre's functionalities, making it easier for users to interact with the process of solving linear systems. It is particularly beneficial for prototyping, offering a quick and efficient means to test various solver and preconditioner configurations. Furthermore, the software allows for the creation of an offline testing framework in which predefined linear systems are read from files and benchmarked with user-defined solution strategies. This makes it an invaluable tool for developers and researchers exploring and validating their computational models. Overall, the software serves as a bridge, bringing the advanced computational capabilities of hypre closer to users who may need more specialized technical expertise, thereby facilitating innovation and exploration in the field of numerical linear algebra.

Paludetto Magri, Victor↗

Porting hypre to heterogeneous computer architectures: Strategies and experiences

We report that linear systems are occurring in many applications, and solving them can take a large amount of the total simulation time. The high performance library hypre provides a variety of interfaces and linear solvers, including various multigrid methods, that have achieved good scalability on a variety of homogeneous parallel computer architectures. Heterogeneous architectures with nodes that have both CPUs and accelerators provide new challenges, since they require more fine-grained parallelism and reduced data movement between different memories on a single node as well as across nodes. We will discuss our experiences and strategies to port hypre to heterogeneous computers with accelerators, including the design of a new memory model, the use of abstractions, the BoxLoop macros in the structured and semi-structured interfaces, and the restructuring of algebraic multigrid (AMG) into modular components. We present numerical experiments comparing CPU and GPU performance for several test problems.

97 MATHEMATICS AND COMPUTING↗

Performance Evaluation of hypre Solvers

This report compares GPU and CPU performance of structured and unstructured hypre solvers on Lassen and Summit. We applied the solvers to various problems of different sizes varying the number of nodes and processes. We provide the individual results for setup phase and solve phase, since preconditioned solvers are often used in different settings. While for a single system solve the total time is determined by the sum of setup and solve time, this can change for different scenarios. For example, in time dependent problems the preconditioner might have to be setup only occasionally or possibly even only once. In that situation solve times will dominate. Thus, since often better GPU speedups can be attained in the solve phase, this will lead to generally better speedups.

97 MATHEMATICS AND COMPUTING↗

Report on hypre performance on AMD GPUs

In this report, the performance of the algebraic multigrid solver used as a preconditioner for conjugate gradient is investigated on 2 nodes of Spock, an early access system at ORNL with 4 MI-100 GPUs and a 64-core Rome CPU per node. We compare GPU and CPU performance for three different diffusion problems using increasing problem sizes.

97 MATHEMATICS AND COMPUTING↗

A New Class of AMG Interpolation Methods Based on Matrix-Matrix Multiplications

A new class of distance-two interpolation methods for algebraic multigrid (AMG) that can be formulated in terms of sparse matrix-matrix multiplications is presented and analyzed. Compared with similar distance-two prolongation operators, the proposed algorithms exhibit improved efficiency and portability to various computing platforms, since they allow one to easily exploit existing high-performance sparse matrix kernels. The new interpolation methods have been implemented in hypre, a widely used parallel multigrid solver library. With the proposed interpolations, the overall time of hypre's BoomerAMG setup can be considerably reduced, while sustaining equivalent, sometimes improved, convergence rates. Numerical results for a variety of test problems on parallel machines are presented that support the superiority of the proposed interpolation operators over the existing ones in hypre.

97 MATHEMATICS AND COMPUTING↗

A New Semistructured Algebraic Multigrid Method

Multigrid methods are well suited to large massively parallel computer architectures because they are mathematically optimal and display good parallelization properties. Since current architecture trends are favoring regular compute patterns to achieve high performance, the ability to express structure has become much more important. The hypre software library provides high-performance multigrid preconditioners and solvers through conceptual interfaces, including a semistructured interface that describes matrices primarily in terms of stencils and logically structured grids. This paper presents a new semistructured algebraic multigrid (SSAMG) method built on this interface. The numerical convergence and performance of a CPU implementation of this method are evaluated for a set of semistructured problems. In conclusion, SSAMG achieves significantly better setup times than hypre’s unstructured AMG solvers and comparable convergence. In addition, the new method is capable of solving more complex problems than hypre’s structured solvers.

97 MATHEMATICS AND COMPUTING↗

Optimizing Irregular Communication with Neighborhood Collectives and Locality-Aware Parallelism

Irregular communication often limits both the performance and scalability of parallel applications. Typically, applications individually implement irregular communication as point-to-point, and any optimizations are integrated directly into the application. As a result, these optimizations lack portability. It is difficult to optimize point-to-point messages within MPI, as the interface for single messages provides no information on the collection of all communication to be performed. However, the persistent neighbor collective API, released in the MPI 4 standard, provides an interface for portable optimizations of irregular communication within MPI libraries. This paper presents methods for implementing existing optimizations for irregular communication within neighborhood collectives, analyzes the impact of replacing point-to-point communication in existing codebases such as Hypre BoomerAMG with neighborhood collectives, and finally shows up to a 1.38x speedup on sparse matrix-vector multiplication communication within a BoomerAMG solve through the use of our optimized neighbor collectives. Here, the authors analyze three implementations of persistent neighborhood collectives for Alltoallv: an unoptimized wrapper of standard point-to-point communication, and two locality-aware aggregating methods. The second locality-aware implementation exposes an non-standard interface to perform additional optimization, and the authors present the additional 0.07x speedup from the extended interface. All optimizations are available in an open-source codebase, MPI Advance, which sits on top of MPI, allowing for optimizations to be added into existing codebases regardless of the system MPI install.

AMG↗

Scaled ILU Smoothers for Navier-Stokes Pressure Projection

Incomplete LU (ILU) smoothers are effective in the algebraic multigrid (AMG) V-cycle for reducing high-frequency components of the error. However, the requisite direct triangular solves are comparatively slow on GPUs. Previous work has demonstrated the advantages of Jacobi iteration as an alternative to direct solution of these systems. Depending on the threshold and fill-level parameters chosen, the factors can be highly nonnormal and Jacobi is unlikely to converge in a low number of iterations. We demonstrate that row scaling can reduce the departure from normality, allowing us to replace the inherently sequential solve with a rapidly converging Richardson iteration. There are several advantages beyond the lower compute time. Scaling is performed locally for a diagonal block of the global matrix because it is applied directly to the factor. Further, an ILUT Schur complement smoother maintains a constant GMRES iteration count as the number of MPI ranks increases, and thus parallel strong-scaling is improved. Our algorithms have been incorporated into hypre, and we demonstrate improved time to solution for linear systems arising in the Nalu-Wind and PeleLM pressure solvers. For large problem sizes, GMRES+AMG executes at least five times faster when using iterative triangular solves compared with direct solves on massively parallel GPUs.

algebraic multigrid↗

AMG2023

The AMG2023 benchmark solves two diffusion problems with a linear solver preconditioned with algebraic multigrid. The code only contains a driver, a Makefile, and a documentation file. It requires an installation of the open source software library hypre that needs to be downloaded elsewhere and is not included here. Its purpose is to benchmark linear solver performance on high performance computers.

Li, Ruipeng↗

AMReX v2024

The software framework, AMReX, supports the development of block-structured adaptive mesh refinement (AMR) algorithms for solving systems of partial differential equations. AMR reduces the computational cost and memory footprint compared to a uniform mesh while preserving the essential local descriptions of different physical processes in complex multiphysics algorithms. AMR uses a hierarchical representation of the solution at multiple levels of resolution where the solution on each level is defined on the union of data containers at that resolution. These data containers, which represent the solution over a logically rectangular subregion of the domain, can contain field data defined on a mesh, Lagrangian particles or combinations of both. In addition to these basic data types, AMReX supports a multilevel embedded boundary representation of complex geometry; linear solvers for cell-centered and nodal data; asynchronous I/O in a native format readable by ParaView, VisIt and yt; and interfaces to hypre and PETSc solvers. AMReX enables applications to run on distributed memory architectures with multicore CPUs and with GPU accelerators. AMReX uses a lightweight abstraction layer that effectively hides the details of the architecture from the application. The framework currently supports CUDA, HIP and SYCL for GPU acceleration and OpenMP for multi-core CPU architectures.

Almgren, Ann↗

ORNL/PFiSM

Simple phase-field model (PFM) based on SAMRAI, Sundials and hypre

Fattebert, Jean-Luc [Oak Ridge National Laboratory↗

Towards Use of Mixed Precision in ECP Math Libraries

The use of multiple types of precision in mathematical software has the potential to increase its performance on new heterogeneous architectures. The xSDK project focuses both on the investigation and development of multiprecision algorithms as well as their inclusion into xSDK member libraries. This report summarizes current efforts on including and/or using mixed precision capabilities in the math libraries Ginkgo, heFFTe, hypre, MAGMA, PETSc/TAO, SLATE, SuperLU, and Trilinos, including KokkosKernels. It contains both numerical results from libraries that already provide mixed precision capabilities, as well as descriptions of the strategies to incorporate multiprecision into established libraries.

97 MATHEMATICS AND COMPUTING↗

Advances in Mixed Precision Algorithms: 2021 Edition

Over the last year, the ECP xSDK-multiprecision effort has made tremendous progress in developing and deploying new mixed precision technology and customizing the algorithms for the hardware deployed in the ECP flagship supercomputers. The effort also has succeeded in creating a cross-laboratory community of scientists interested in mixed precision technology and now working together in deploying this technology for ECP applications. In this report, we highlight some of the most promising and impactful achievements of the last year. Among the highlights we present are: Mixed precision IR using a dense LU factorization and achieving a 1.8× speedup on Spock; results and strategies for mixed precision IR using a sparse LU factorization; a mixed precision eigenvalue solver; Mixed Precision GMRES-IR being deployed in Trilinos, and achieving a speedup of 1.4× over standard GMRES; compressed Basis (CB) GMRES being deployed in Ginkgo and achieving an average 1.4× speedup over standard GMRES; preparing hypre for mixed precision execution; mixed precision sparse approximate inverse preconditioners achieving an average speedup of 1.2×; and detailed description of the memory accessor separating the arithmetic precision from the memory precision, and enabling memory-bound low precision BLAS 1/2 operations to increase the accuracy by using high precision in the computations without degrading the performance. We emphasize that many of the highlights presented here have also been submitted to peer-reviewed journals or established conferences, and are under peer-review or have already been published.

97 MATHEMATICS AND COMPUTING↗

Advances in Mixed Precision Algorithms: 2021 Edition

Over the last year, the ECP xSDK-multiprecision effort has made tremendous progress in developing and deploying new mixed precision technology and customizing the algorithms for the hardware deployed in the ECP flagship supercomputers. The effort also has succeeded in creating a cross-laboratory community of scientists interested in mixed precision technology and now working together in deploying this technology for ECP applications. In this report, we highlight some of the most promising and impactful achievements of the last year. Among the highlights we present are • Mixed precision IR using a dense LU factorization and achieving a 1.8× speedup on Spock; • Results and strategies for mixed precision IR using a sparse LU factorization; • A mixed precision eigenvalue solver; • Mixed Precision GMRES-IR being deployed in Trilinos, and achieving a speedup of 1.4× over standard GMRES; • Compressed Basis (CB) GMRES being deployed in Ginkgo and achieving an average 1.4× speedup over standard GMRES; • Preparing hypre for mixed precision execution; • Mixed precision sparse approximate inverse preconditioners achieving an average speedup of 1.2×; • Detailed description of the memory accessor separating the arithmetic precision from the memory precision, and enabling memory-bound low precision BLAS 1/2 operations to increase the accuracy by using high precision in the computations without degrading the performance. We emphasize that many of the highlights presented here have also been submitted to peer-reviewed journals or established conferences, and are under peer-review or have already been published.

97 MATHEMATICS AND COMPUTING↗

FY2022 Q4: Demonstrate multi-turbine simulation with hybrid-structured/unstructured-moving-grid software stack running primarily on GPUs and propose improvements for successful KPP-2 [Poster]

Milestone accomplishments staged the ExaWind team for successful completion of KPP-2 challenge problem in FY23, which requires the simulation on Frontier of at least four MW-scale turbines in an atmospheric boundary layer with at least 20B gridpoints. The ExaWind project and software stack is many faceted, with team members working on multiple areas, including linear-system solvers (Trilinos, hypre, AMReX), overset meshes, turbulence modeling, and in situ visualization, all with an aim for high fidelity predictions and performance portability. This milestone marks significant improvements on many fronts and provides the team with a pathway to exascale wind farm simulations in FY23.

17 WIND ENERGY↗

Batched Sparse Linear Algebra (Final Report for Subcontract B648960)

This report finalizes design specifications for developing batched kernels for small tensor operations for unassembled matrix-free iterative solvers, batched solvers for partially assembled operators, and batched solvers with support for various sparse formats. The outcome of the project milestones is a set of interfaces to Batched Sparse LA solvers running on hardware accelerators for use in ECP Libraries and Applications. It is part of the development of sparse batched kernels, solvers/preconditioners as well as creating interoperability in xSDK libraries with sparse and dense batched functions to benefit ECP applications. The participants included representatives from ECP libraries (not limited to the xSDK project), applications, and vendors (AMD, Intel, and NVIDIA). Batched sparse linear algebra solvers form the new frontier for algorithmic development and performance engineering. Many applications (ECP and non-ECP alike) require simultaneous solutions of small linear systems of equations that are structurally sparse. To move towards high hardware utilization, it is important to provide these applications with appropriate interfaces to efficient batched sparse solvers running on modern hardware accelerators. We present interface designs in use by HPC software libraries supporting batched sparse linear algebra and the development of sparse batched kernel codes for solvers and preconditioners. We also address the potential interoperability opportunities to keep the software portable between the major hardware accelerators from AMD, Intel, and NVIDIA. The presented interface specifications includes batched band, sparse iterative, and sparse direct solvers. This report summarizes progress in Kokkos Kernels and the xSDK libraries MAGMA, Ginkgo, hypre, SUNDIALS, and SuperLU_dist.

97 MATHEMATICS AND COMPUTING↗