Engineering Papers⌕ Search

Engineering topics

Carns, Philip

Publications and source records attributed to Carns, Philip.

Performance Characterization and Provenance of Distributed Task-based Workflows on HPC Platforms

Understanding performance and provenance of task-based workflows poses significant challenges, particularly in distributed configurations where resources are shared by multiple applications. Task-based workflow management systems further complicate performance predictability because of their dynamicity that subtly alters task execution order from run to run. In this paper we propose a layered characterization framework for performance and task provenance for Dask.distributed workflows running on high-performance computing (HPC) platforms. It collects data from jobs, the workflow management system, and the operating system to aid in understanding the performance of these workflows. Our approach encompasses three main contributions: first, an extension of Dask.distributed to capture high-fidelity task provenance using Mochi data services; second, the adaptation of the established HPC I/O characterization tool Darshan to gather high-fidelity I/O data, thereby enhancing the granularity of our analysis; and third, a framework to combine and process the collected data and provide helpful insights into performance characterization and reproducibility, alongside our lessons learned.

Dask↗

Autonomy Loops for Monitoring, Operational Data Analytics, Feedback, and Response in HPC Operations

Many High Performance Computing (HPC) facilities have developed and deployed frameworks in support of continuous monitoring and operational data analytics (MODA) to help improve efficiency and throughput. Because of the complexity and scale of systems and workflows and the need for low-latency response to address dynamic circumstances, automated feedback and response have the potential to be more effective than current human-in-the-loop approaches which are laborious and error prone. Progress has been limited, however, by factors such as the lack of infrastructure and feedback hooks, and successful deployment is often site- and case-specific. In this position paper we report on the outcomes and plans from a recent Dagstuhl Seminar, seeking to carve a path for community progress in the development of autonomous feedback loops for MODA, based on the established formalism of similar (MAPE-K) loops in autonomous computing and self-adaptive systems. By defining and developing such loops for significant cases experienced across HPC sites, we seek to extract commonalities and develop conventions that will facilitate interoperability and interchangeability with system hardware, software, and applications across different sites, and will motivate vendors and others to provide telemetry interfaces and feedback hooks to enable community development and pervasive deployment of MODA autonomy loops.

autonomy loops↗

HEPnOS: a Specialized Data Service for High Energy Physics Analysis

In this paper, we present HEPnOS, a distributed data service for managing data produced by high-energy physics (HEP) experiments. Using HEPnOS, HEP applications can use HPC resources more effciently than traditional fle-based applications. The fle-based model leads to a rigid, chunk-based allocation of computational resources and limits the number of cores that can be used concurrently by an HEP application. The fundamental problem is that organizing domain-specifc data into fles inadvertently introduces a single, artifcial, confated tuning parameter that puts key optimization goals into confict: larger fle sizes reduce metadata overhead and thus improve I/O effciency, but smaller fle sizes provide more opportunity for workfow parallelism and load balancing. In this work, we introduce a domain-specifc data service that decouples that constraint so that data can be accessed and processed in its natural granularity while still maintaining I/O effciency. By removing the constraints introduced by fle handling we are able to obtain better scaling and make effcient use of more cores for processing a fxed-sized data sample. We demonstrate the improved scalability by using an application developed in the fle-based paradigm and comparing it to a version modifed to use HEPnOS.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗