Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Many-core”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

31 records · Page 2

Exascale models of stellar explosions: Quintessential multi-physics simulation

The ExaStar project aims to deliver an efficient, versatile, and portable software ecosystem for multi-physics astrophysics simulations run on exascale machines. The code suite is a component-based multi-physics toolkit, built on the capabilities of current simulation codes (in particular Flash-X and Castro), and based on the massively parallel adaptive mesh refinement framework AMReX. It includes modules for hydrodynamics, advanced radiation transport, thermonuclear kinetics, and nuclear microphysics. The code will reach exascale efficiency by building upon current multi- and many-core packages integrated into an orchestration system that uses a combination of configuration tools, code translators, and a domain-specific asynchronous runtime to manage performance across a range of platform architectures. The target science includes multi-physics simulations of astrophysical explosions (such as supernovae and neutron star mergers) to understand the cosmic origin of the elements and the fundamental physics of matter and neutrinos under extreme conditions.

79 ASTRONOMY AND ASTROPHYSICS↗

Enhancing scalability of a matrix-free eigensolver for studying many-body localization

We propose several techniques to enhance the parallel scalability of a matrix-free eigensolver designed for studying many-body localization (MBL) of quantum spin chain models with nearest-neighbor interactions and on-site disorder. This type of problem is computationally challenging because the dimension of the associated Hamiltonian matrix grows exponentially with respect to the number of spins L, and we need to average over different realizations of the random disorder to obtain relevant statistical behavior. For each disorder realization, we need to compute eigenvalues from different regions of the spectrum and their corresponding eigenvectors. In previous work, the interior eigenstates for a single eigenvalue problem are computed via the shift-and-invert Lanczos algorithm. Due to the extremely high memory footprint of the LU factorizations, this technique is not well suited for large L’s. For example, we need thousands of compute nodes on modern high performance computing infrastructures to go beyond L = 24. The matrix-free approach does not suffer from this memory bottleneck, however, its scalability is limited by a computation and communication load imbalance. To reduce this imbalance and to significantly enhance the scalability of the matrix-free eigensolver, we reorder the matrix and leverage the consistent space runtime, CSPACER. We also show its efficiency in managing irregular communication patterns at scale compared to optimized MPI non-blocking two-sided and one-sided RMA implementation variants. This effort enables us to study MBL for spin chains with a larger number of spins. The efficiency and effectiveness of the proposed algorithm is demonstrated by computing eigenstates on a massively parallel many-core high performance computer.

METIS↗

Fixing Amdahl's Law within the Limits of Accelerated Systems: FALLACY

Closeout report for FALLACY project. The performance of Data Model Convergence Initiative (DMC) applications on parallel machines is far below the limit set by Amdahl’s law. Whether the machine is based on many-core, GPUs, FPGAs, or a heterogeneous combination, usually the most significant bottleneck is accessing data from the memory system. Aligning with DMC’s HW/architecture thrust, this project developed a set of memory-centric tools called ‘MemGaze’ that inform the HW/SW stack about an application’s memory behavior, including data access latency and diagnosing poor data layout and data composition. Our approach uses architectural modeling and analysis of workload data accesses.

97 MATHEMATICS AND COMPUTING↗

Green's function methods for multiphysics simulations (Final Report)

Green's functions are important tools for analyzing mathematical properties of partial differential equations (PDEs), and for numerically solving PDEs, especially when equations for the same operator but with multiple right hand sides need to be solved simultaneously. This proposal aims at developing efficient and accurate numerical methods for computing Green's functions, which can be used to tackle a challenging question in multiphysics simulation of DOE-mission science problems: how to couple quantum physics with classical physics. The key mathematical difficulty is properly formulate a "boundary condition" for the region described by quantum physics, and conventional approaches often model such boundary conditions in an empirical way. The proposed Green's function methods use the Dirichlet-to-Neumann map to formulate a boundary condition that is non-empirical and can couple the quantum and classical regions in an in principle exact way. The key ingredient of the new methods is to construct the Dirichlet-to-Neumann map in an efficient, accurate and versatile manner. The new methods have provably low complexity and are ideally suited for massively parallel and emerging many-core computational systems.

97 MATHEMATICS AND COMPUTING↗

Graph Contractions for Calculating Correlation Functions in Lattice QCD

Computing correlation functions for many-particle systems in Lattice QCD is vital to extract nuclear physics observables like the energy spectrum of hadrons such as protons. However, this type of calculation has long been considered to be very challenging and computing-resource intensive because of the complex nature of a hadron composed of quarks with many degrees of freedom. In particular, a correlation function can be calculated through a sum of all possible pairs of quark contractions, each of which is a batched tensor contraction, dictated by Wick's theorem. Because the number of terms of this sum can be very large for any hadronic system of interest, fast evaluation of the sum faces several challenges: an extremely large number of contractions, a huge memory footprint at runtime, and the speed of tensor contractions. In this paper, we present a Lattice QCD analysis software suite, Redstar, which addresses these challenges by utilizing novel algorithmic and software engineering methods targeting modern computing platforms such as many-core CPUs and GPUs. In particular, Redstar represents every term in the sum of a correlation function by a graph, applies efficient graph algorithms to reduce the number of contractions to lower the cost of computations, and minimizes the total memory footprint. Moreover, Redstar carries out the contractions on either CPUs or GPUs utilizing an internal and highly efficient Hadron contraction library. Specifically, we illustrate some important algorithmic optimizations of Redstar, show various key design features of Hadron library, and present the speedup values due to the optimizations along with performance figures for calculating six correlations functions on four computing platforms.

Chen, Jie↗

Technical descriptions of the experimental dynamical downscaling simulations over North America by the CAM–MPAS variable-resolution model

Abstract. Comprehensive assessment of climate datasets is important for communicating model projections and associated uncertainties to stakeholders. Uncertainties can arise not only from assumptions and biases within the model but also from external factors such as computational constraint and data processing. To understand sources of uncertainties in global variable-resolution (VR) dynamical downscaling, we produced a regional climate dataset using the Model for Prediction Across Scales (MPAS; dynamical core version 4.0) coupled to the Community Atmosphere Model (CAM; version 5.4), which we refer to as CAM–MPAS hereafter. This document provides technical details of the model configuration, simulations, computational requirements, post-processing, and data archive of the experimental CAM–MPAS downscaling data. The CAM–MPAS model is configured with VR meshes featuring higher resolutions over North America as well as quasi-uniform-resolution meshes across the globe. The dataset includes multiple uniform- (240 and 120 km) and variable-resolution (50–200, 25–100, and 12–46 km) simulations for both the present-day (1990–2010) and future (2080–2100) periods, closely following the protocol of the North American Coordinated Regional Climate Downscaling Experiment. A deviation from the protocol is the pseudo-warming experiment for the future period, using the ocean boundary conditions produced by adding the sea surface temperature and sea-ice changes from the low-resolution version of the Max Planck Institute Earth System Model (MPI-ESM-LR) in the Coupled Model Intercomparison Project Phase 5 to the present-day ocean state from a reanalysis product. Some unique aspects of global VR models are evaluated to provide background knowledge to data users and to explore good practices for modelers who use VR models for regional downscaling. In the coarse-resolution domain, strong resolution sensitivity of the hydrological cycles exists over the tropics but does not appear to affect the midlatitude circulations in the Northern Hemisphere, including the downscaling target of North America. The pseudo-warming experiment leads to similar responses of large-scale circulations to the imposed radiative and boundary forcings in the CAM–MPAS and MPI-ESM-LR models, but their climatological states in the historical period differ over various regions, including North America. Such differences are carried to the future period, suggesting the importance of the base state climatology. Within the refined domain, precipitation statistics improve with higher resolutions, and such statistical inference is verified to be negligibly influenced by horizontal remapping during post-processing. Limited (≈50 % slower) throughput of the current code is found on a recent many-core/wide-vector high-performance computing system, which limits the lengths of the 12–46 km simulations and indirectly affects sampling uncertainty. Our experience shows that global and technical aspects of the VR downscaling framework require further investigations to reduce uncertainties for regional climate projection.

58 GEOSCIENCES↗

Toward GEOS-6, A Global Cloud System Resolving Atmospheric Model

NASA is committed to observing and understanding the weather and climate of our home planet through the use of multi-scale modeling systems and space-based observations. Global climate models have evolved to take advantage of the influx of multi- and many-core computing technologies and the availability of large clusters of multi-core microprocessors. GEOS-6 is a next-generation cloud system resolving atmospheric model that will place NASA at the forefront of scientific exploration of our atmosphere and climate. Model simulations with GEOS-6 will produce a realistic representation of our atmosphere on the scale of typical satellite observations, bringing a visual comprehension of model results to a new level among the climate enthusiasts. In preparation for GEOS-6, the agency's flagship Earth System Modeling Framework [JDl] has been enhanced to support cutting-edge high-resolution global climate and weather simulations. Improvements include a cubed-sphere grid that exposes parallelism; a non-hydrostatic finite volume dynamical core, and algorithm designed for co-processor technologies, among others. GEOS-6 represents a fundamental advancement in the capability of global Earth system models. The ability to directly compare global simulations at the resolution of spaceborne satellite images will lead to algorithm improvements and better utilization of space-based observations within the GOES data assimilation system

Putman, William M.↗

GPU Accelerated Vector Median Filter

Noise reduction is an important step for most image processing tasks. For three channel color images, a widely used technique is vector median filter in which color values of pixels are treated as 3-component vectors. Vector median filters are computationally expensive; for a window size of n x n, each of the n(sup 2) vectors has to be compared with other n(sup 2) - 1 vectors in distances. General purpose computation on graphics processing units (GPUs) is the paradigm of utilizing high-performance many-core GPU architectures for computation tasks that are normally handled by CPUs. In this work. NVIDIA's Compute Unified Device Architecture (CUDA) paradigm is used to accelerate vector median filtering. which has to the best of our knowledge never been done before. The performance of GPU accelerated vector median filter is compared to that of the CPU and MPI-based versions for different image and window sizes, Initial findings of the study showed 100x improvement of performance of vector median filter implementation on GPUs over CPU implementations and further speed-up is expected after more extensive optimizations of the GPU algorithm .

Aras, Rifat↗

Accelerating Climate and Weather Simulations through Hybrid Computing

Unconventional multi- and many-core processors (e.g. IBM (R) Cell B.E.(TM) and NVIDIA (R) GPU) have emerged as effective accelerators in trial climate and weather simulations. Yet these climate and weather models typically run on parallel computers with conventional processors (e.g. Intel, AMD, and IBM) using Message Passing Interface. To address challenges involved in efficiently and easily connecting accelerators to parallel computers, we investigated using IBM's Dynamic Application Virtualization (TM) (IBM DAV) software in a prototype hybrid computing system with representative climate and weather model components. The hybrid system comprises two Intel blades and two IBM QS22 Cell B.E. blades, connected with both InfiniBand(R) (IB) and 1-Gigabit Ethernet. The system significantly accelerates a solar radiation model component by offloading compute-intensive calculations to the Cell blades. Systematic tests show that IBM DAV can seamlessly offload compute-intensive calculations from Intel blades to Cell B.E. blades in a scalable, load-balanced manner. However, noticeable communication overhead was observed, mainly due to IP over the IB protocol. Full utilization of IB Sockets Direct Protocol and the lower latency production version of IBM DAV will reduce this overhead.

hybrid computing↗

m-CUBES An efficient and portable implementation of multi-dimensional integration for gpus

The task of multi-dimensional numerical integration is frequently encountered in physics and other scientific fields, e.g., in modeling the effects of systematic uncertainties in physical systems and in Bayesian parameter estimation. Multi-dimensional integration is often time-prohibitive on CPUs. Efficient implementation on many-core architectures is challenging as the workload across the integration space cannot be predicted a priori. We propose m-Cubes, a novel implementation of the well-known Vegas algorithm for execution on GPUs. Vegas transforms integration variables followed by calculation of a Monte Carlo integral estimate using adaptive partitioning of the resulting space. m-Cubes improves performance on GPUs by maintaining relatively uniform workload across the processors. As a result, our optimized Cuda implementation for Nvidia GPUs outperforms parallelization approaches proposed in past literature. We further demonstrate the efficiency of m-Cubes by evaluating a six-dimensional integral from a cosmology application, achieving significant speedup and greater precision than the CUBA library's CPU implementation of VEGAS. We also evaluate m-Cubes on a standard integrand test suite. m-Cubes outperforms the serial implementations of the Cuba and GSL libraries by orders of magnitude speedup while maintaining comparable accuracy. Our approach yields a speedup of at least 10 when compared against publicly available Monte Carlo based GPU implementations. In summary, m-Cubes can solve integrals that are prohibitively expensive using standard libraries and custom implementations. A modern C++ interface header-only implementation makes m-Cubes portable, allowing its utilization in complicated pipelines with easy to define stateful integrals. Compatibility with non-Nvidia GPUs is achieved with our initial implementation of m-Cubes using the Kokkos framework.

Sakiotis, Ioannis↗

A compute-bound formulation of Galerkin model reduction for linear time-invariant dynamical systems

This work aims to advance computational methods for projection-based reduced-order models (ROMs) of linear time-invariant (LTI) dynamical systems. For such systems, current practice relies on ROM formulations expressing the state as a rank-1 tensor (i.e., a vector), leading to computational kernels that are memory bandwidth bound and, therefore, ill-suited for scalable performance on modern architectures. This weakness can be particularly limiting when tackling many-query studies, where one needs to run a large number of simulations. This work introduces a reformulation, called rank-2 Galerkin, of the Galerkin ROM for LTI dynamical systems which converts the nature of the ROM problem from memory bandwidth to compute bound. We present the details of the formulation and its implementation, and demonstrate its utility through numerical experiments using, as a test case, the simulation of elastic seismic shear waves in an axisymmetric domain. We quantify and analyze performance and scaling results for varying numbers of threads and problem sizes. In conclusion, we present an end-to-end demonstration of using the rank-2 Galerkin ROM for a Monte Carlo sampling study. We show that the rank-2 Galerkin ROM is one order of magnitude more efficient than the rank-1 Galerkin ROM (the current practice) and about 970 times more efficient than the full-order model, while maintaining accuracy in both the mean and statistics of the field.

97 MATHEMATICS AND COMPUTING↗

Toward performance-portable PETSc for GPU-based exascale systems

The Portable Extensible Toolkit for Scientific computation (PETSc) library delivers scalable solvers for nonlinear time-dependent differential and algebraic equations and for numerical optimization. The PETSc design for performance portability addresses fundamental GPU accelerator challenges and stresses flexibility and extensibility by separating the programming model used by the application from that used by the library, and it enables application developers to use their preferred programming model, such as Kokkos, RAJA, SYCL, HIP, CUDA, or OpenCL, on upcoming exascale systems. Furthermore, a blueprint for using GPUs from PETSc-based codes is provided, and case studies emphasize the flexibility and high performance achieved on current GPU-based systems.

97 MATHEMATICS AND COMPUTING↗

Low-synch Gram–Schmidt with delayed reorthogonalization for Krylov solvers

The parallel strong-scaling of iterative methods is often determined by the number of global reductions at each iteration. Low-synch Gram-Schmidt algorithms are applied here to the Arnoldi algorithm to reduce the number of global reductions and therefore to improve the parallel strong-scaling of iterative solvers for nonsymmetric matrices such as the GMRES and the Krylov-Schur iterative methods. In the Arnoldi context, the factorization is "left-looking" and processes one column at a time. Among the methods for generating an orthogonal basis for the Arnoldi algorithm, the classical Gram-Schmidt algorithm, with reorthogonalization (CGS2) requires three global reductions per iteration. A new variant of CGS2 that requires only one reduction per iteration is presented and applied to the Arnoldi algorithm. Delayed CGS2 (DCGS2) employs the minimum number of global reductions per iteration (one) for a one-column at-a-time algorithm. The main idea behind the new algorithm is to group global reductions by rearranging the order of operations. DCGS2 must be carefully integrated into an Arnoldi expansion or a GMRES solver. Numerical stability experiments assess robustness for Krylov-Schur eigenvalue computations. Performance experiments on the ORNL Summit supercomputer then establish the superiority of DCGS2 over CGS2.

97 MATHEMATICS AND COMPUTING↗