Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Autotuning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

27 records · Page 2

HPC Network Simulation Tuning via Automatic Extraction of Hardware Parameters

Popular HPC network interconnection simulators such as SST/macro provide a variety of configurable parameters to explore the design space of hardware components such as network interface cards (NIC), switches, and links among them. While such knobs provide flexibility to explore design trade-offs for novel hardware, manually configuring simulations for matching configurations of the existing hardware to focus on topology exploration can be cumbersome and error-prone, leading to widely inaccurate simulations. This challenge is compounded when specifications of various (proprietary) technologies are not readily available or intentionally omitted. In this work, we propose a framework to autotune the multiple network models’ simulation configurations within SST/macro using Tree-structured Parzen Estimator-based Bayesian optimization to observe the effect on simulation accuracy across different message regimes. These regimes consist of small to large message sizes and latency to bandwidth-bound messages. We provide a detailed analysis of the simulation error for four representative HPC systems. Our Bayesian optimization based autotuning framework for network models achieves a maximum of 5x improvement in accuracy over best-effort manual configurations based on available hardware specifications.

Simulation, autotuning↗

Hydrogen maser frequency standard computer model for automatic cavity tuning servo simulations

A computer model of the JPL hydrogen maser frequency standard was developed. This model allows frequency stability data to be generated, as a function of various maser parameters, many orders of magnitude faster than these data can be obtained by experimental test. In particular, the maser performance as a function of the various automatic tuning servo parameters may be readily determined. Areas of discussion include noise sources, first-order autotuner loop, second-order autotuner loop, and a comparison of the loops.

Potter, P. D.↗

Cross-Feature Transfer Learning for Efficient Tensor Program Generation

Tuning tensor program generation involves navigating a vast search space to find optimal program transformations and measurements for a program on the target hardware. The complexity of this process is further amplified by the exponential combinations of transformations, especially in heterogeneous environments. This research addresses these challenges by introducing a novel approach that learns the joint neural network and hardware features space, facilitating knowledge transfer to new, unseen target hardware. A comprehensive analysis is conducted on the existing state-of-the-art dataset, TenSet, including a thorough examination of test split strategies and the proposal of methodologies for dataset pruning. Leveraging an attention-inspired technique, we tailor the tuning of tensor programs to embed both neural network and hardware-specific features. Notably, our approach substantially reduces the dataset size by up to 53% compared to the baseline without compromising Pairwise Comparison Accuracy (PCA). Furthermore, our proposed methodology demonstrates competitive or improved mean inference times with only 25–40% of the baseline tuning time across various networks and target hardware. The attention-based tuner can effectively utilize schedules learned from previous hardware program measurements to optimize tensor program tuning on previously unseen hardware, achieving a top-5 accuracy exceeding 90%. This research introduces a significant advancement in autotuning tensor program generation, addressing the complexities associated with heterogeneous environments and showcasing promising results regarding efficiency and accuracy.

97 MATHEMATICS AND COMPUTING↗

Modeling pre-Exascale AMR Parallel I/O Workloads via Proxy Applications

The present work investigates the modeling of preexascale input/output (I/O) workloads of Adaptive Mesh Refinement (AMR) simulations through a simple proxy application. We collect data from the AMReX Castro framework running on the Summit supercomputer for a wide range of scales and mesh partitions for the hydrodynamic Sedov case as a baseline to provide sufficient coverage to the formulated proxy model. The non-linear analysis data production rates are quantified as a function of a set of input parameters such as output frequency, grid size, number of levels, and the Courant-Friedrichs-Lewy (CFL) condition number for each rank, mesh level and simulation time step. Linear regression is then applied to formulate a simple analytical model which allows to translate AMReX inputs into MACSio proxy I/O application parameters, resulting in a simple “kernel” approximation for data production at each time step. Results show that MACSio can simulate actual AMReX nonlinear “static” I/O workloads to a certain degree of confidence on the Summit supercomputer using the present methodology. The goal is to provide an initial level of understanding of AMR I/O workloads via lightweight proxy applications models to facilitate autotune data management strategies in anticipation of exascale systems.

Godoy, William↗

A MultiGPU Performance-Portable Solution for Array Programming Based on Kokkos

Today, multiGPU nodes are widely used in high-performance computing and data centers. However, current programming models do not provide simple, transparent, and portable support for automatically targeting multiple GPUs within a node on application areas of array programming. In this paper, we describe a new application programming interface based on the Kokkos programming model to enable array computation on multiple GPUs in a transparent and portable way across both NVIDIA and AMD GPUs. We implement different variations of this technique to accommodate the exchange of stencils (array boundaries) among different GPU memory spaces, and we provide autotuning to select the proper number of GPUs, depending on the computational cost of the operations to be computed on arrays, that is completely transparent to the programmer. We evaluate our multiGPU extension on Summit (#5 TOP500), with six NVIDIA V100 Volta GPUs per node, and Crusher that contains identical hardware/software as Frontier (#1 TOP500), with four AMD MI250X GPUs, each with 2 Graphics Compute Dies (GCDs)for a total of 8 GCDs per node. We also compare the performance of this solution against the use of MPI + Kokkos, which is the cur-rent de facto solution for multiple GPUs in Kokkos. Our evaluation shows that the new Kokkos solution provides good scalability for many GPUs and a faster and simpler solution (from a programming productivity perspective) than MPI + Kokkos.

Valero Lara, Pedro↗

NASA atomic hydrogen standards program - An update

Some of the design features of NASA hydrogen masers are discussed including the large hydrogen source bulb, the palladium purified, the state selector, the replaceable pumps, the small entrance stem, magnetic shields, the elongated storage bulb, the aluminum cavity, the electronics package, and the autotuner. Attention is also given to the reliability and operating life of these hydrogen atomic standards.

Reinhardt, V. S.↗

Frequency, phase, and amplitude changes of the hydrogen maser oscillation

The frequency, the phase, and the amplitude changes of the hydrogen maser oscillation, which are induced by the modulation of the cavity resonant frequency, are considered. The results obtained apply specifically to one of the H-maser cavity autotuning methods which is actually implemented, namely the cavity frequency-switching method. The frequency, the phase, and the amplitude changes are analyzed theoretically. The phase and the amplitude variations are measured experimentally. It is shown, in particular, that the phase of oscillation is subjected to abrupt jumps at the times of the cavity frequency switching, whose magnitude is specified. The results given can be used for the design of a phase-locked loop (PLL) aimed at minimizing the transfer of the phase modulation to the slaved VCXO.

Audoin, Claude↗

Matrix Multiply Performance of GPUs on Exascale-class HPE/Cray Systems

The computation of dense matrix-matrix products (GEMMs) is central to many modeling and simulation workloads as well as AI/ML deep learning campaigns. In fact, millions of dollars are spent annually on computing GEMMs, and large model training demands are increasing exponentially. Specialized processors such as GPUs are designed to perform well for these operations. However, the performance of GEMMs on GPUs can exhibit complex behaviors depending on many factors, making it challenging to optimize the performance of GEMMs on these processors. In this study we undertake an examination of GEMM performance on several leading GPU models taken from product lines of GPUs to be deployed in forthcoming exascale computing systems. We show results to illustrate the many factors that can affect performance of GEMMs on GPUs. We then present data collected from a large number of test runs for an example GEMM operation to show the dependence behaviors of GEMM rate on matrix dimensions. Finally, we show results from machine learning-based performance models using novel feature engineering methods to fit the measured performance, providing a potential basis for GEMM performance tuning and autotuning methods for GPUs. Recommendations are also given for how to achieve high GEMM performance on modern GPUs.

Melesse Vergara, Veronica↗