Predictive Modeling and Diagnostic Monitoring of Extreme Science Workflows (Final Report)
This proposal addresses a critical issue of performance prediction identified in the report from the ASCR \Computational Modeling of Big Networks (COMBINE)" workshop: "end-to-end performance is not predictable due to a variety of factors. Even when some performance forecasts or predictions can be made, they often cannot explain the reasons why some predictions fail." We will develop new analytical models to predict the end-to-end performance of scientific workflows on DOE computing infrastructures, and use simulations and experimentation to validate and refine these models, as well as to pinpoint the sources of model inaccuracy. We will also use these models to help diagnose application and infrastructure problems, and to adapt the system based on this diagnosis. This section provides background in the areas relevant to the proposed work. RPI’s specific tasks within the Panorama project are as follows: (1) Develop Aspen-Simulation interface for Workflow Model Driven Simulation. (2) Validate manual performance models of two target workflow scenarios with empirical measurement and simulation. (3) Extend ROSS-Aspen API to simulate workflow descriptions when required. (4) Validate Aspen performance models of two target workflow scenarios with automatic performance model empirical measurement and simulation. (5) Design and implement final system to automatically generate Aspen performance models from workflow descriptions (including methods to compensate for limitations of Aspen analytical models). (6) Validate improved Aspen performance models with target workflow on production infrastructure. To date, all the project milestones assigned to us where reached within the best of our abilities over the course of the project performance period. Below describes the key outcome from our collaborative research in a system named, Durango .