Engineering PapersSearch

DOE OSTI · 3739679

Fine-Grained Power and Energy Attribution on AMD GPU/APU-Based Exascale Nodes

Abstract

Modern exascale GPU- and APU-based systems provide multiple power and energy sensors, but differences in scope, update rate, timing, and filtering complicate the attribution of short-lived accelerator activity. This paper presents a methodology to characterize and correct these effects on Cray EX systems with AMD Instinct MI250X GPUs (Frontier) and MI300A APUs (Portage). Using controlled square-wave workloads, we quantify update intervals, delay, aliasing, and variability across up to 512 GPUs and 480 APUs with on-chip (rocm-smi/amd-smi) and off-chip Cray Power Management sensors. We reconstruct power from cumulative energy counters to achieve faster response times, validate it against on-chip, off-chip, and node-level sensors, and integrate the resulting streams into a Score-P/PAPI-based tool for time-aligned, phase-level attribution. Applied to rocHPL, rocHPL-MxP, and HPG-MxP, the method separates energy savings due to reduced runtime from changes in power. Mixed precision reduces node energy on Frontier by 79% for rocHPL-MxP and 31% for HPG-MxP, with similar trends on Portage. These results provide portable guidance for sensor validation and power-aware optimization on current and future exascale systems.

Keep this discovery

BibTeXRIS

Mcdaniel, Adam [ORNL] (ORCID:000000016926028X), Jantz, Michael R. [University of Tennessee, Knoxville (UTK)], Sharma, Ashesh [Hewlett Packard Enterprise], Martin, Steven [Hewlett Packard Enterprise], Abbott, Steve [Hewlett Packard Enterprise], Khandekar, Shreyas [HPE], Neth, Brandon [Hewlett Packard Enterprise], Alvarez, Bruno Villasenor [Advanced Micro Devices (AMD)], Kashi, Aditya [ORNL] (ORCID:0000000325893792), Elwasif, Wael [ORNL] (ORCID:0000000305541036), Hernandez Mendoza, Oscar [ORNL] (ORCID:0000000253806951). 2026-06-01. Fine-Grained Power and Energy Attribution on AMD GPU/APU-Based Exascale Nodes. https://doi.org/10.23919/isc.2026.11520492

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related reports

Vidyut3d: A Gpu Accelerated Fluid Solver for Non-Equilibrium Plasmas on Adaptive Grids

We present the numerical methods, programming methodology, verification, and performance assessment of a non-equilibrium plasma fluid solver that can effectively utilize current and upcoming central processing and graphics processing unit (CPU+GPU) architectures, in this work. Our plasma fluid model solves the coupled conservation equations for species transport, electrostatic Poisson and electron temperature on adaptive Cartesian grids. Our solver is written using performance portable adaptive-grid/particle management library, AMReX, and is portable over widely available vendor specific GPU architectures. We present verification of our solver using method of manufactured solutions that indicate formal second order accuracy with central diffusion and fifth-order weighted-essentially-non-oscillatory (WENO) advection scheme. We also verify our solver with published literature on capacitive discharges and atmospheric pressure streamer propagation. We demonstrate the use of our solver on two 3D simulation cases: an atmospheric streamer propagation in Ar-H2 mixtures and a low pressure twin electrode radio frequency reactor. Our performance studies on three different CPU+GPU architectures indicate approximately 150-400X speed-up using AMD and NVIDIA GPUs per time step compared to a single CPU core for a 4 million cell simulation with 15 species.

Sitaraman, Hariswaran

Establishing model credibility for process-microstructure-property relationships in additive manufacturing using exascale computing

Additive Manufacturing (AM) of alloys holds significant promise as a disruptive technology in various industries, yet its adoption is often hindered by challenges in achieving consistent part quality. These issues are primarily due to the complex process-microstructure-property (PSP) relationships inherent to AM. Computational models can greatly aid in understanding these relationships, but their widespread impact and adoption has been limited by a lack of validated, open-source, and computationally efficient PSP modeling frameworks and hardware limitations. Here, this study leverages the ExaAM software suite and data from the AMBench-2018 series of laser powder bed fusion (LPBF) benchmark experiments to perform a comprehensive model assessment, including verification, validation, sensitivity analysis, and uncertainty quantification. The RADICAL-EnTK workflow manager was used to perform an ensemble of heat transport, solidification, and mechanical response simulations on the exascale computer Frontier, considering uncertainties in critical model inputs such as laser spot size and nucleation parameters, and consisting of 125 explicit grain structure simulations and 7875 crystal plasticity simulations. For a selected location within the Inconel 625 AMBench-2018 test artifact, sensitivity analysis and uncertainty quantification were performed using the predicted distributions of grain structure and mechanical properties. Qualitative agreement was found between the predicted grain size and texture and the observed AMBench-2018 microstructure, the mean predicted yield stress was within 5% of the experimental measurement mean, and the mean predicted engineering stress at 5% strain was within 10% of the experimental measurement mean. The insights gained from development and validation of the ExaAM PSP modeling framework will help guide future directions for enhancing the credibility and reliability of PSP models in AM, thereby accelerating the adoption of AM technologies in various industries.

Additive manufacturing

Dust Composition of Comet 81P/Wild 2 From JWST Spectroscopy Compared to Stardust’s Fine-Grained Materials and Gems-Rich IDPS

The Stardust Mission returned samples from the coma of Jupiter Family comet 81P/Wild 2 for detailed laboratory analyses. Here we present and discuss the best-fit thermal dust model for the coma dust of comet 81P/Wild 2 as observed by JWST. Comet 81P was observed through JWST GO 018xx, using NIRSpec (2.9–5.3µm, λ/∆λ≈1000) and MRS IFU (4.9–28.1µm, λ/∆λ≈3000) on 2023-03-20 UT and 2023-03-24 UT, respectively, at a heliocentric distance of 1.85 au and JWST distance of 1.43 au (phase angle of 32 degrees). The dust coma of 81P as revealed by JWST offers a salient compliment to laboratory studies of Stardust samples that typically are bigger than 2 µm and up to 60 µm in size.

D H Wooden