Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “GESTS”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

GPU-enabled extreme-scale turbulence simulations: Fourier pseudo-spectral algorithms at the exascale using OpenMP offloading

Fourier pseudo-spectral methods for nonlinear partial differential equations are of wide interest in many areas of advanced computational science, including direct numerical simulation of three-dimensional (3-D) turbulence governed by the Navier-Stokes equations in fluid dynamics. This paper presents a new capability for simulating turbulence at a new record resolution up to 35 trillion grid points, on the world's first exascale computer, Frontier, comprising AMD MI250x GPUs with HPE's Slingshot interconnect and operated by the US Department of Energy's Oak Ridge Leadership Computing Facility (OLCF). Key programming strategies designed to take maximum advantage of the machine architecture involve performing almost all computations on the GPU which has the same memory capacity as the CPU, performing all-to-all communication among sets of parallel processes directly on the GPU, and targeting GPUs efficiently using OpenMP offloading for intensive number-crunching including 1-D Fast Fourier Transforms (FFT) performed using AMD ROCm library calls. With 99% of computing power on Frontier being on the GPU, leaving the CPU idle leads to a net performance gain via avoiding the overhead of data movement between host and device except when needed for some I/O purposes. Memory footprint including the size of communication buffers for MPI_ALLTOALL is managed carefully to maximize the largest problem size possible for a given node count. Detailed performance data including separate contributions from different categories of operations to the elapsed wall time per step are reported for five grid resolutions, from 2048 3 on a single node to 32768 3 on 4096 or 8192 nodes out of 9408 on the system. Both 1D and 2D domain decompositions which divide a 3D periodic domain into slabs and pencils respectively are implemented. The present code suite (labeled by the acronym GESTS, GPUs for Extreme Scale Turbulence Simulations) achieves a figure of merit (in grid points per second) exceeding goals set in the Center for Accelerated Application Readiness (CAAR) program for Frontier. The performance attained is highly favorable in both weak scaling and strong scaling, with notable departures only for 2048 3 where communication is entirely intra-node, and for 32768 3 , where a challenge due to small message sizes does arise. Communication performance is addressed further using a lightweight test code that performs all-to-all communication in a manner matching the full turbulence simulation code. Performance at large problem sizes is affected by both small message size due to high node counts as well as dragonfly network topology features on the machine, but is consistent with official expectations of sustained performance on Frontier. Overall, although not perfect, the scalability achieved at the extreme problem size of 32768 3 (and up to 8192 nodes — which corresponds to hardware rated at just under 1 exaflop/sec of theoretical peak computational performance) is arguably better than the scalability observed using prior state-of-the-art algorithms on Frontier's predecessor machine (Summit) at OLCF. New science results for the study of intermittency in turbulence enabled by this code and its extensions are to be reported separately in the near future.

3D fast Fourier transform↗

Outcomes of OpenMP Hackathon: OpenMP Application Experiences with the Offloading Model (Part II)

This paper reports on experiences gained and practices adopted when using the latest features of OpenMP to port a variety of HPC applications and mini-apps based on different computational motifs (BerkeleyGW, WDMApp/XGC, GAMESS, GESTS, and GridMini) to accelerator-based, leadership-class, high-performance supercomputer systems at the Department of Energy. As recent enhancements to OpenMP become available in implementations, there is a need to share the results of experimentation with them in order to better understand their behavior in practice, to identify pitfalls, and to learn how they can be effectively deployed in scientific codes. Additionally, we identify best practices from these experiences that we can share with the rest of the OpenMP community.

Chapman, Barbara↗

Outcomes of OpenMP Hackathon: OpenMP Application Experiences with the Offloading Model (Part I)

This paper reports on experiences gained and practices adopted when using the latest features of OpenMP to port a variety of HPC applications and mini-apps based on different computational motifs (BerkeleyGW, WDMApp/XGC, GAMESS, GESTS, and GridMini) to accelerator-based, leadership-class, high-performance supercomputer systems at the Department of Energy. As recent enhancements to OpenMP become available in implementations, there is a need to share the results of experimentation with them in order to better understand their behavior in practice, to identify pitfalls, and to learn how they can be effectively deployed in scientific codes. Additionally, we identify best practices from these experiences that we can share with the rest of the OpenMP community.

Chapman, Barbara↗

OpenMP application experiences: Porting to accelerated nodes

As recent enhancements to the OpenMP specification become available in its implementations, there is a need to share the results of experimentation in order to better understand the OpenMP implementation’s behavior in practice, to identify pitfalls, and to learn how the implementations can be effectively deployed in scientific codes. We report on experiences gained and practices adopted when using OpenMP to port a variety of ECP applications, mini-apps and libraries based on different computational motifs to accelerator-based leadership-class high-performance supercomputer systems at the United States Department of Energy. Additionally, we identify important challenges and open problems related to the deployment of OpenMP. Through our report of experiences, we find that OpenMP implementations are successful on current supercomputing platforms and that OpenMP is a promising programming model to use for applications to be run on emerging and future platforms with accelerated nodes.

97 MATHEMATICS AND COMPUTING↗

A red-emitting carborhodamine for monitoring and measuring membrane potential

Biological membrane potentials, or voltages, are a central facet of cellular life. Optical methods to visualize cellular membrane voltages with fluorescent indicators are an attractive complement to traditional electrode-based approaches, since imaging methods can be high throughput, less invasive, and provide more spatial resolution than electrodes. Recently developed fluorescent indicators for voltage largely report changes in membrane voltage by monitoring voltage-dependent fluctuations in fluorescence intensity. However, it would be useful to be able to not only monitor changes but also measure values of membrane potentials. This study discloses a fluorescent indicator which can address both. We describe the synthesis of a sulfonated tetramethyl carborhodamine fluorophore. When this carborhodamine is conjugated with an electron-rich, methoxy (-OMe) containing phenylenevinylene molecular wire, the resulting molecule, CRhOMe, is a voltage-sensitive fluorophore with red/far-red fluorescence. Using CRhOMe, changes in cellular membrane potential can be read out using fluorescence intensity or lifetime. In fluorescence intensity mode, CRhOMe tracks fast-spiking neuronal action potentials (APs) with greater signal-to-noise than state-of-the-art BeRST 1 (another voltage-sensitive fluorophore). CRhOMe can also measure values of membrane potential. The fluorescence lifetime of CRhOMe follows a single exponential decay, substantially improving the quantification of membrane potential values using fluorescence lifetime imaging microscopy (FLIM). The combination of red-shifted excitation and emission, mono-exponential decay, and high voltage sensitivity enable fast FLIM recording of APs in cardiomyocytes. The ability to both monitor and measure membrane potentials with red light using CRhOMe makes it an important approach for studying biological voltages.

59 BASIC BIOLOGICAL SCIENCES↗