Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Performance Portability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Users manual for the Automated Performance Test System (APTS)

The characteristics of and the user information for the Essex Automated Performance Test System (APTS) computer-based portable performance assessment battery are given. The battery was developed to provide a menu of performance test tapping the widest possible variety of human cognitive and motor functions, implemented on a portable computer system suitable for use in both laboratory and field settings for studying the effects of toxic agents and other stressors. The manual gives guidance in selecting, administering and scoring tests from the battery, and reviews the data and studies underlying the development of the battery. Its main emphasis is on the users of the battery - the scientists, researchers and technicians who wish to examine changes in human performance across time or as a function of changes in the conditions under which test data are obtained. First the how to information needed to make decisions about where and how to use the battery is given, followed by the research background supporting the battery development. Further, the development history of the battery focuses largely on the logical framework within which tests were evaluated.

Lane, N. E.↗

A tool and a methodology to use macros for abstracting variations in code for different computational demands

Scientific software used on high-performance computing platforms is in a phase of transformation because of the combined increase in the heterogeneity and complexity of models and hardware platforms. Having separate implementations for different platforms can easily lead to combinatorial explosions; therefore, the computational science community has been looking for mechanisms to express code through abstractions that can be specialized for different platforms. Most existing approaches use template meta-programming in C++, and are, therefore language specific. Here, we have developed a tool that uses customized expansion of macros to mimic some of C++ behavior in other languages. It enables unification of any code variants that may be necessary to run efficiently on different target architectures and different computational environments through use of macros with multiple alternative definitions and ability to arbitrate on definition selection for expansion. Combined with two other tools, a custom runtime, and a user specified recipe translator, our custom macroprocessor becomes a part of an overall performance portability solution that does not depend on any specific programming language. We also use macros as code-shorthand that lets code snippets become building blocks that allow variations in control flow to explore performance options. We demonstrate use of macros in Flash-X, a multiphysics multicomponent code with many Fortran legacy components derived from an earlier community code FLASH.

Heterogenous computing↗

bolt: A GPU-Accelerated Solver for Marginally Collisional Kinetic Physics Using a High- Resolution Constrained Transport Scheme

Collisionless and marginally collisional plasmas are very expensive to simulate due to high dimensionality and large scale separations, especially in global problems with additional length scales. bolt addresses this challenge by achieving high performance on GPUs using the arrayfire performance portability library. We have demonstrated accuracy on a range of test problems, and the bolt code paper is currently in preparation.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

From PeleC to PeleACC, to PeleC++

PeleC is an Exascale Computing Project application for simulating compressible combustion in complex geometries. It has been built on top of the popular AMReX library. In the beginning of the Exascale Computing Project, PeleC was focused on KNL. It uses a mixture of C++, C, and kernels written in Fortran to obtain performance by focusing on vectorization. Recently we have taken two approaches in deciding PeleC's future for obtaining performance on exascale GPU machines. In the first programming model, we decorated the Fortran kernels with OpenACC directives. This expedited our ability to run at large scales on Summit's GPUs, where we achieved a significant speedup over the CPUs on Summit. The second programming model involved rewriting the Fortran kernels in C++ and using AMReX's Kokkos-like lambda abstractions for running on the GPU. This resulted in similar speedups on Summit's GPUs over merely utilizing the CPUs. Both approaches involved AMReX's management of memory transfers between the device and host. In this work, we compare and contrast the benefits and pitfalls to both programming approaches regarding performance, performance portability, and productivity. We also discuss advantages we have found in taking the time to modernize our code and why have chosen a specific pathway to prepare our code for the future DOE exascale machines.

exascale computing↗

Vidyut3d: A GPU accelerated fluid solver for non-equilibrium plasmas on adaptive grids

We present the numerical methods, programming methodology, verification, and performance assessment of a non-equilibrium plasma fluid solver that can effectively utilize current and upcoming central processing and graphics processing unit (CPU+GPU) architectures, in this work. Our plasma fluid model solves the coupled conservation equations for species transport, electrostatic Poisson and electron temperature on adaptive Cartesian grids. Our solver is written using performance portable adaptive-grid/particle management library, AMReX, and is portable over widely available vendor specific GPU architectures. We present verification of our solver using method of manufactured solutions that indicate formal second order accuracy with central diffusion and fifth-order weighted-essentially-non-oscillatory (WENO) advection scheme. We also verify our solver with published literature on capacitive discharges and atmospheric pressure streamer propagation. We demonstrate the use of our solver on two 3D simulation cases: an atmospheric streamer propagation in Ar-H2 mixtures and a low pressure three-electrode radio frequency reactor. Our performance studies on three different CPU+GPU architectures indicate ~ 150-400X speed-up using AMD and NVIDIA GPUs per time step compared to a single CPU core for a 4 million cell simulation with 15 species.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

MAGMA: Enabling exascale performance with accelerated BLAS and LAPACK for diverse GPU architectures

MAGMA (Matrix Algebra for GPU and Multicore Architectures) is a pivotal open-source library in the landscape of GPU-enabled dense and sparse linear algebra computations. With a repertoire of approximately 750 numerical routines across four precisions, MAGMA is deeply ingrained in the DOE software stack, playing a crucial role in high-performance computing. Notable projects such as ExaConstit, HiOP, MARBL, and STRUMPACK, among others, directly harness the capabilities of MAGMA. In addition, the MAGMA development team has been acknowledged multiple times for contributing to the vendors’ numerical software stacks. Looking back over the time of the Exascale Computing Project (ECP), we highlight how MAGMA has adapted to recent changes in modern HPC systems, especially the growing gap between CPU and GPU compute capabilities, as well as the introduction of low precision arithmetic in modern GPUs. We also describe MAGMA’s direct impact on several ECP projects. Maintaining portable performance across NVIDIA and AMD GPUs, and with current efforts toward supporting Intel GPUs, MAGMA ensures its adaptability and relevance in the ever-evolving landscape of GPU architectures.

97 MATHEMATICS AND COMPUTING↗

kokkos-fft: A shared-memory FFT for the Kokkos ecosystem

kokkos-fft provides a unified, performance-portable interface for Fast Fourier Transforms (FFTs) within the Kokkos ecosystem (C. Trott et al., 2021). It seamlessly integrates with leading local FFT libraries including FFTW, cuFFT, rocFFT, and oneMKL. Designed for simplicity and efficiency, kokkos-fft offers a user experience akin to numpy.fft for in-place and out-of-place transforms, while leveraging the raw speed of vendor-optimized libraries. A demonstration solving 2D Hasegawa-Wakatani turbulence with the Fourier spectral method illustrates how kokkos-fft can deliver significant speedups over Python-based alternatives without drastically increasing code complexity, empowering researchers to perform high-performance FFTs simply and effectively.

97 MATHEMATICS AND COMPUTING↗

Flash-X: A multiphysics simulation software instrument

Flash-X is a highly composable multiphysics software system that can be used to simulate physical phenomena in several scientific domains. It derives some of its solvers from FLASH, which was first released in 2000. Flash-X has a new framework that relies on abstractions and asynchronous communications for performance portability across a range of increasingly heterogeneous hardware platforms. Flash-X is meant primarily for solving Eulerian formulations of applications with compressible and/or incompressible reactive flows. It also has a built-in, versatile Lagrangian framework that can be used in many different ways, including implementing tracers, particle-in-cell simulations, and immersed boundary methods.

97 MATHEMATICS AND COMPUTING↗

Trilinos: Enabling Scientific Computing across Diverse Hardware Architectures at Scale

Trilinos is a community-developed, open-source software framework that facilitates building large-scale, complex, multiscale, multiphysics simulation code bases for scientific and engineering problems. Since the Trilinos framework has undergone substantial changes to support new applications and new hardware architectures, this document is an update to “An Overview of the Trilinos project” by Heroux et al. (ACM Transactions on Mathematical Software, 31(3):397–423, 2005). It describes the design of Trilinos, introduces its new organization in product areas, and highlights established and new features available in Trilinos. Particular focus is put on the modernized software stack based on the Kokkos ecosystem to deliver performance portability across heterogeneous hardware architectures. This article also outlines the organization of the Trilinos community and the contribution model to help onboard interested users and contributors.

Heterogeneous Hardware Architectures↗

Application of Portable Parallelization Strategies for GPUs on track reconstruction kernels

Utilizing the computational power of GPUs is one of the key ingredients to meet the computing challenges presented to the next generation of High-Energy Physics (HEP) experiments. Unlike CPUs, developing software for GPUs often involves using architecturespecific programming languages promoted by the GPU vendors and hence limits the platform that the code can run on. Various portability solutions have been developed to achieve portable, performant software across different GPU vendors. Given the rapid evolution of these portability solutions, an early adoption of them in simple HEP testbed applications will help us understand the strengths and weaknesses of respective approaches.We apply several portability solutions, including Alpaka, Kokkos, SYCL and std::execution::par, on kernels for track propagation extracted from the mkFit project. We report on the development experience of the same application with different portability solutions, as well as their performance on GPUs, measured as the throughput of the kernels, from different manufacturers such as NVIDIA, AMD and Intel.

Kwok, Martin [Fermilab] (ORCID:0000000286936146)↗

A Sparse Distributed Gigascale Resolution Material Point Method

In this paper, we present a four-layer distributed simulation system and its adaptation to the Material Point Method (MPM). The system is built upon a performance portable C++ programming model targeting major High-Performance-Computing (HPC) platforms. A key ingredient of our system is a hierarchical block-tile-cell sparse grid data structure that is distributable to an arbitrary number of Message Passing Interface (MPI) ranks. We additionally propose strategies for efficient dynamic load balance optimization to maximize the efficiency of MPI tasks. Our simulation pipeline can easily switch among backend programming models, including OpenMP and CUDA, and can be effortlessly dispatched onto supercomputers and the cloud. Finally, we construct benchmark experiments and ablation studies on supercomputers and consumer workstations in a local network to evaluate the scalability and load balancing criteria. We demonstrate massively parallel, highly scalable, and gigascale resolution MPM simulations of up to 1.01 billion particles for less than 323.25 seconds per frame with 8 OpenSSH-connected workstations.

97 MATHEMATICS AND COMPUTING↗

ORNL/grit

Grit is a software library designed specifically for conducting particle-based Lagrangian simulations with CPU/GPUperformance portability. This library allows researchers and engineers to perform a wide range of simulations, including multiphase flow, and more. Grit employs Message Passing Interface (MPI) for distributed memory parallelism and Kokkos programming model for on-node shared memory parallelism with performance portability across different architectures of GPUs and multi-core/manycore CPUs.

Ge, Wenjun [Oak Ridge National Laboratory (ORNL), ↗

Lessons Learned and Scalability Achieved When Porting Uintah to DOE Exascale Systems

A key challenge faced when preparing codes for Department of Energy (DOE) exascale systems was designing scalable applications for systems featuring hardware and software not yet available at leadership-class scale. With such systems now available, it is important to evaluate scalability of the resulting software solutions on these target systems. One such code designed with the exascale DOE Aurora and DOE Frontier systems in mind is the Uintah Computational Framework, an open-source asynchronous many-task (AMT) runtime system. To prepare for exascale, Uintah adopted a portable MPI+X hybrid parallelism approach using the Kokkos performance portability library (i.e., MPI+Kokkos). This paper complements recent work with additional details and an evaluation of the resulting approach on Aurora and Frontier. Results are shown for a challenging benchmark demonstrating interoperability of 3 portable codes essential to Uintah-related combustion research. These results demonstrate single-source portability across Aurora and Frontier with scaling characteristics shown to 3,072 Aurora nodes and 9,216 Frontier nodes. In addition to showing results run to new scales on new systems, this paper also discusses lessons learned through efforts preparing Uintah for exascale systems.

Holmen, John [ORNL] (ORCID:0000000259342641)↗

The Rodinia Benchmark Suite in SYCL

As opposed to the Open Computing Language (OpenCL) programming model in which host and device codes are generally written in different programming languages, SYCL can combine host and device codes for an application in a type-safe way to improve development productivity and performance portability. Rodinia is a widely used benchmark suite for heterogeneous computing. Hence, we port the OpenCL implementations of the benchmark suite to SYCL manually. The SYCL benchmark suite is an opensource project for tracking the development of the mainstream SYCL compilers, and for developers and researchers interested in programming productivity, performance analysis, and portability across different computing platforms. We organize the remainder of the report as follows. Section II introduces the SYCL programming model, shows the major differences between an OpenCL program and a SYCL program, and gives a summary of the benchmark suite. In Section III, we describe our SYCL implementations in more details. We evaluate the SYCL benchmarks on the Intel® central processing units (CPUs) and graphics processing units (GPUs) in Section IV. Section V concludes the report.

97 MATHEMATICS AND COMPUTING↗

Discrete-Element and Material-Point Method (DEM and MPM) Based Solvers for Sustainable Technologies

We present the use of discrete element method (DEM) and material point method (MPM) in three relevant green technology applications that include biomass feedstock handling, lithium-ion battery manufacturing, and high-pressure reverse osmosis. Our open-source DEM and MPM solvers are developed using performance portable grid and particle management library, AMReX, thus enabling superior performance on NVIDIA and AMD GPUs with > 100 million particles. Our DEM solver resolves the motion of individual particles in a granular system and includes a bonded sphere method for modeling non-spherical particles along with Hertzian and liquid bridge-based contact models. We simulate highly variable biomass feedstock flows in large-scale hoppers for biofuel production and electrode calendering in battery manufacturing using DEM. Our simulations predict flow blockage in large scale biomass hoppers and electrode microstructure variations, thus providing valuable information for biofuel and battery manufacturers, respectively. The second half of the talk will be on MPM and its application towards pore resolved simulations of reverse osmosis membranes under compressive loads. We present a validation study of our MPM simulations with membrane microscopy imaging thus providing useful insights on membrane stability under high pressure conditions. We also present a spectral stability analysis of using linear hat, quadratic and cubic spline basis in MPM indicating regions of numerical stability.

BIOMASS FUELS,MATHEMATICS AND COMPUTING↗

Climbing the Summit and Pushing the Frontier of Mixed Precision Benchmarks at Extreme Scale

The rise of machine learning (ML) applications and their use of mixed precision to perform interesting science are driving forces behind AI for science on HPC. The convergence of ML and HPC with mixed precision offers the possibility of transformational changes in computational science. The HPL-AI benchmark is designed to measure the performance of mixed precision arithmetic as opposed to the HPL benchmark which measures double precision performance. Pushing the limits of systems at extreme scale is nontrivial -little public literature explores optimization of mixed precision computations at this scale. In this work, we demonstrate how to scale up the HPL-AI benchmark on the pre-exascale Summit and exascale Frontier systems at the Oak Ridge Leadership Computing Facility (OLCF) with a cross-platform design. We present the implementation, performance results, and a guideline of optimization strategies employed for delivering portable performance on both AMD and NVIDIA GPUs at extreme scale.

Lu, Hao↗

Leveraging Compiler-Based Translation to Evaluate a Diversity of Exascale Platforms

Accelerator-based heterogeneous computing is the de facto standard in current and upcoming exascale machines. These heterogeneous resources empower computational scientists to select a machine or platform well-suited to their domain or applications. However, this diversity of machines also poses challenges related to programming model selection: inconsistent availability of programming models across different exascale systems, lack of performance portability for those programming models that do span several systems, and inconsistent performance between different models on a single platform. We explore these challenges on exascale-similar hardware, including AMD MI100 and NVIDIA A100 GPUs. By extending the sourceto-source compiler OpenARC, we demonstrate the power of automated translation of applications written in a single frontend programming model (OpenACC) into a variety of backend models (OpenMP, OpenCL, CUDA, HIP) that span the upcoming exascale environments. This translation enables us to compare performance within and across devices and to analyze programming model behavior with profiling tools.

Lambert, Jacob↗

XaaS: Acceleration as a Service to Enable Productive High-Performance Cloud Computing

High-performance computing (HPC) and the cloud have evolved independently, specializing their innovations into performance or productivity. Acceleration as a Service (XaaS) is a recipe to empower both fields with a shared execution platform that provides transparent access to computing resources, regardless of the underlying cloud or HPC service provider. Bridging HPC and cloud advancements, XaaS presents a unified architecture built on performance-portable containers. Here, our converged model concentrates on low-overhead, high-performance communication and computing, targeting resource-intensive workloads from climate simulations to machine learning. XaaS lifts the restricted allocation model of Function as a Service (FaaS), allowing users to benefit from the flexibility and efficient resource utilization of serverless computing while supporting long-running and performance-sensitive workloads from HPC.

97 MATHEMATICS AND COMPUTING↗