Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Workflow Management Systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Seamless integration of commercial Clouds with ATLAS Distributed Computing

The CERN ATLAS Experiment successfully uses a worldwide dis-tributed computing Grid infrastructure to support its physics programme at the Large Hadron Collider (LHC). The Grid workflow system PanDA routinely manages up to 700,000 concurrently running production and analysis jobs to process simulation and detector data. In total more than 500 PB of data are distributed over more than 150 sites in the WLCG and handled by the ATLAS data management system Rucio. To prepare for the ever growing data rate in future LHC runs new developments are underway to embrace industry accepted protocols and technologies, and utilize opportunistic resources in a standard way. This paper reviews how the Google and Amazon Cloud computing ser-vices have been seamlessly integrated as a Grid site within PanDA and Rucio. Performance and brief cost evaluations will be discussed. Such setups could offer advanced Cloud tool-sets and provide added value for analysis facilities that are under discussions for LHC Run-4.

97 MATHEMATICS AND COMPUTING↗

Operational experience and R&D results using the Google Cloud for High-Energy Physics in the ATLAS experiment

The ATLAS experiment at CERN relies on a Worldwide Distributed Computing Grid infrastructure to support its physics program at the Large Hadron Collider. ATLAS has integrated cloud computing resources to complement its Grid infrastructure and conducted an R&D program on Google Cloud Platform. These initiatives leverage key features of commercial cloud providers: lightweight configuration and operation, elasticity and availability of diverse infrastructures. Here this paper examines the seamless integration of cloud computing services as a conventional Grid site within the ATLAS workflow management and data management systems, while also offering new setups for interactive, parallel analysis. It underscores pivotal results that enhance the on-site computing model and outlines several R&D projects that have benefited from large-scale, elastic resource provisioning models. Furthermore, this study discusses the impact of cloud-enabled R&D projects in three domains: accelerators and AI/ML, ARM CPUs and columnar data analysis techniques.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Position Papers for the ASCR Workshop on the Management and Storage of Scientific Data

The purpose of this workshop is to identify priority research directions in the area of data management for high-performance and scientific computing above and beyond HPC’s traditional "the parallel file system is the data-management system" model. Supporting the breadth of the DOE mission, including the explosion of AI uses and the growing needs of experimental and observational science, motivates revisiting our assumptions about data management. There are many facets of this topic to explore including: (1) Interfaces for accessing data that resides on traditional persistent storage as well as memory devices; (2) Storage-system architecture design that supports scientific workflows on varied hierarchical storage and networking devices; (3) Devising metadata management infrastructure to support FAIR principles (Findability, Accessibility, Interoperability, and Reusability); (4) Capturing provenance information about scientific data; (5) Utilizing AI to learn I/O patterns of emerging workloads for efficient data management; (6) Providing data management support for AI and complex workflows; and (7) Understanding the overlap between traditional storage systems and I/O (SSIO) efforts and data management. While the program committee has identified these topics as important areas for discussion, we welcome position papers from the community that propose additional topics of interest for discussion at the workshop. The workshop agenda will include breakout sessions for discussing these and selected topic areas to inform priority research directions for data management for high-performance and scientific computing.

97 MATHEMATICS AND COMPUTING↗

Report for the ASCR Workshop on the Management and Storage of Scientific Data

The purpose of this workshop is to identify priority research directions in the area of data management for high-performance and scientific computing above and beyond HPC’s traditional "the parallel file system is the data-management system" model. Supporting the breadth of the DOE mission, including the explosion of AI uses and the growing needs of experimental and observational science, motivates revisiting our assumptions about data management. There are many facets of this topic to explore including: (1) Interfaces for accessing data that resides on traditional persistent storage as well as memory devices; (2) Storage-system architecture design that supports scientific workflows on varied hierarchical storage and networking devices; (3) Devising metadata management infrastructure to support FAIR principles (Findability, Accessibility, Interoperability, and Reusability); (4) Capturing provenance information about scientific data; (5) Utilizing AI to learn I/O patterns of emerging workloads for efficient data management; (6) Providing data management support for AI and complex workflows; and (7) Understanding the overlap between traditional storage systems and I/O (SSIO) efforts and data management. While the program committee has identified these topics as important areas for discussion, we welcome position papers from the community that propose additional topics of interest for discussion at the workshop. The workshop agenda will include breakout sessions for discussing these and selected topic areas to inform priority research directions for data management for high-performance and scientific computing.

97 MATHEMATICS AND COMPUTING↗

Pipeline for Integrated Projects in Energy Systems (PIPES): A Tool for Integrated System Planning [Slides]

The Pipeline for Integrated Projects in Energy Systems (PIPES) is a comprehensive project, data, and workflow management tool designed for integrated modeling teams. PIPES facilitates the management of data requirements, tasks, and progress tracking, serving as a higher-level integration layer that works across various data and modeling software. This tool integrates models, data, and tools to perform large-scale, integrated analysis work at scale. PIPES is designed to streamline integrated modeling projects, enhance collaboration, and ensure the quality and efficiency of data management and workflow processes. This presentation introduces PIPES a multi-model tool for integrated system planning; it describes the underlying architecture, deep dives into common user workflows, and outlines the upcoming development roadmap beyond its current alpha state.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

CAMEO: A Co-design Architecture for Multi-objective Energy System Optimization (Project Report)

CAMEO (Codesign Architecture for Multi-objective Energy System Optimization) is a modular workflow management framework that abstracts co-design problems as Directed Acyclic Graphs (DAG). The framework employs JSON-based workflow specifications that enable systematic decomposition of complex optimization problems into reusable, interchangeable components including data loaders, scenario generators, optimization solvers, and result summarizers.

97 MATHEMATICS AND COMPUTING↗

Optimization of distributed compute resources utilization in the CMS Global Pool

The CMS Submission Infrastructure is the primary system for managing computing resources for CMS workflows, including data processing, simulation, and analysis. It integrates geographically distributed resources from Grid, HPC, and cloud providers into federated pools managed by HTCondor and Glidein- WMS, for a total of around 500k CPU cores. This system dynamically manages workloads based on priorities defined by the collaboration. Additionally, CMS scheduling strategies must be flexible to handle multiple concurrent workloads while considering changing processing demands and resource availability from various providers.Efficient utilization of vast amounts of distributed compute resources is a key element for the success of the scientific programs of the LHC experiments. Optimizing the system is essential to maximize resource efficiency and fully utilize the distributed computing power. The CMS Submission Infrastructure team thus systematically investigates sources of inefficiency in workload scheduling to reduce their impact. In addition, a strategy of pilot overloading has been introduced to compensate for other inefficiency sources, thereby optimizing resource utilization and enhancing computational throughput.

Mascheroni, Marco [UC, San Diego (main)]↗

Automated and Distributed Monte Carlo Generation for GlueX

MCwrapper is a set of systems that manages the entire Monte Carlo production workflow for GlueX and provides standards for how that Monte Carlo is produced. MCwrapper was designed to be able to utilize a variety of batch systems in a way that is relatively transparent to the user, thus enabling users to quickly and easily produce valid simulated data at home institutions worldwide. Additionally, MCwrapper supports an autonomous system that takes user’s project submissions via a custom web application. The system then atomizes the project into individual jobs, matches these jobs to resources, and monitors the jobs status. The entire system is managed by a database which tracks almost all facets of the systems from user submissions to the individual jobs themselves. Users can interact with their submitted projects online via a dashboard or, in the case of testing failure, can modify their project requests from a link contained in an automated email. Beginning in 2018 the GlueX Collaboration began to utilize the Open Science Grid (OSG) to handle a bulk of simulation tasks; these tasks are currently being performed on the OSG automatically via MCwrapper. This talk will outline the entire system of MCwrapper, its use cases, and the unique challenges facing the system.

Britton, Thomas↗

Integrated End-to-end Performance Prediction and Diagnosis for Extreme Scientific Workflows

This report details recent progress for the ASCR funded project “Integrated End-to-end Performance Prediction and Diagnosis for Extreme Scientific Workflows”. We refer to the project as IPPD/2, reflecting the 2017 renewal under expanded scope and partners In IPPD/2, we increased our research scope to include data motion. We are focusing on three major aspects: a) observe how data is generated, distributed, and used; b) analyze how data is (repeatedly) consumed with a focus both on repeated patterns and anomalies; and c) explore how to optimize data motion. This new work on data motion will augment and complement IPPD/2’s research that focused on the computational aspects of tasks. We leverage and extend our existing tools and demonstrate our work on the Belle II workflow suite as well as on workflows from NSLS-II. The highlights of our work are as follows: Provenance for Workflows: Provenance is used to provide information enabling quality control, re-run computational workflows, and reproduce results. IPPD/2 has been building a scalable provenance management system that enables the capture of provenance from the high-level workflow through all relevant system levels in one integrated environment. Leveraging this work, our recent efforts have included using provenance as an enabling technique. Workload characterization: Leveraging provenance and analysis, we characterize data movement within network, storage, and memory over a variety of workloads. This characterization enables an understanding by performance analysts and application developers of the range of behaviors that could be expected. Performance Prediction for Workflows: The goal of modeling distributed workflows is to understand performance bottlenecks and enable more intelligent task scheduling to optimize selected metrics of interest (e.g., task throughput or output data rate). IPPD/2 has utilized both analytical and AI/ML modeling methodologies for performance modeling. Advanced Scheduling and Fault Modeling for Workflows: Scheduling of large-scale scientific workflows on geographically distributed resources is a challenging problem. To improve workflow throughput, we combined novel scheduling algorithms with task predictions from performance modeling and fault modeling. Dynamically Alleviating Bottlenecks in Workflows: Exploiting our provenance, analysis, and modeling efforts, we have explored and developed several techniques for dynamically detecting and alleviating bottlenecks in data movement. In particular, we have spent considerable effort demonstrating our techniques on production-like workflow configurations.

97 MATHEMATICS AND COMPUTING↗

RISE: Reducing I/O Contention in Staging-based Extreme-Scale In-situ Workflows

While in-situ workflow formulations have addressed some of the data-related challenges associated with extreme-scale scientific workflows, these workflows involve complex interactions and different modes of data exchange. In the context of increasing system complexity, such workflows present significant resource management challenges, requiring complex cost-performance tradeoffs. This paper presents RISE, an intelligent staging-based data management middleware, which builds on the DataSpaces framework and performs intelligent scheduling of data management operations to reduce I/O contention. In RISE, data are always written immediately to local buffers to reduce the effect of the transfer impact upon application performance. RISE identifies applications’ data access patterns and moves data towards data consumers only when the network is expected to be idle, reducing the impact of asynchronous background data movement upon critical data read/write requests. Here, we experimentally demonstrate that RISE can take advantage of staging nodes to offload data during writes without degrading application data movement performance.

97 MATHEMATICS AND COMPUTING↗

MADA: Multi-Agent Design Assistant

MADA (Multi-Agent Design Assistant) is a Large Language Model (LLM) powered multi-agent framework that coordinates specialized agents for complex design workflows. The system was designed for HPC workflows with the following agents in mind: 1) A Job Management Agent (JMA) launches and manages ensemble simulations on HPC systems, 2) a Geometry Agent (GA) generates meshes, and 3) an Inverse Design Agent (IDA) proposes new designs informed by simulation outcomes. Our framework reduces cumbersome manual workflow setup, and enables automated design exploration at scale. However, the software also enables users to rapidly create new multi-agent systems. Simply define new agents in a configuration file, giving each their own set of tools (via MCP), and then chat and prompt your new multi-agent system. Is

Gunnarson, BrianS [Lawrence Livermore National Lab↗

First-generation college student forges ahead, now key to Lab's mission

Fatima Woody was just 17 and a student at Pojoaque Valley High School when she first started her career as a Los Alamos Neutron Science Center receptionist. Now she's in a crucial role that keeps plutonium pit production and other mission processes operating with as little interruption as possible. Over nearly four decades, Fatima has gradually advanced from her initial positions as a receptionist and administrative secretary to become a computer technician, then a computer system professional who specializes in project management. Today, she coordinates the workflow for nearly two dozen deployed information technology technicians who keep the computer systems and networks operating at the high-tech complex that houses the Lab's Plutonium Facility.

99 GENERAL AND MISCELLANEOUS↗

Software and computing for Run 3 of the ATLAS experiment at the LHC

The ATLAS experiment has developed extensive software and distributed computing systems for Run 3 of the LHC. These systems are described in detail, including software infrastructure and workflows, distributed data and workload management, database infrastructure, and validation. The use of these systems to prepare the data for physics analysis and assess its quality are described, along with the software tools used for data analysis itself. An outlook for the development of these projects towards Run 4 is also provided.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

$\mathrm{RADICAL}$-Pilot and $\mathrm{PMIx}$/$\mathrm{PRRTE}$: Executing Heterogeneous Workloads at Large Scale on Partitioned $\mathrm{HPC}$ Resources

Execution of heterogeneous workflows on high-performance computing (HPC) platforms present unprecedented resource management and execution coordination challenges for runtime systems. Task heterogeneity increases the complexity of resource and execution management, limiting the scalability and efficiency of workflow execution. Re-source partitioning and distribution of tasks execution over portioned re-sources promises to address those problems but we lack an experimental evaluation of its performance at scale. Here this paper provides a performance evaluation of the Process Management Interface for Exascale (PMIx) and its reference implementation PRRTE on the leadership-class HPC plat-form Summit, when integrated into a pilot-based runtime system called RADICAL-Pilot. We partition resources across multiple PRRTE Distributed Virtual Machine (DVM) environments, responsible for launching tasks via the PMIx interface. We experimentally measure the work-load execution performance in terms of task scheduling/launching rate and distribution of DVM task placement times, DVM startup and termination overheads on the Summit leadership-class HPC platform. Integrated solution with PMIx/PRRTE enables using an abstracted, standardized set of interfaces for orchestrating the launch process, dynamic process management and monitoring capabilities. It extends scaling capabilities allowing to overcome a limitation of other launching mechanisms (e.g., JSM/LSF). Explored different DVM setup configurations provide insights on DVM performance and a layout to leverage it. Our experimental results show that heterogeneous workload of 65,500 tasks on 2048 nodes, and partitioned across 32 DVMs, runs steady with resource utilization not lower than 52%. While having less concurrently executed tasks resource utilization is able to reach up to 85%, based on results of heterogeneous workload of 8200 tasks on 256 nodes and 2 DVMs.

97 MATHEMATICS AND COMPUTING↗

Empowering Scientific Discovery Through Computing at the Advanced Photon Source

This paper explores the challenges and solutions for managing and processing the vast amount of data generated by the Advanced Photon Source (APS), a synchrotron light source facility producing ultra-bright x-rays for diverse scientific domains. With 68 experimental beamlines covering materials research, biology, and more, the APS serves a wide user base across academia, government, and industry. The ongoing upgrade of the APS storage ring and installation of new instruments will amplify data generation and processing demands. This paper discusses the approach to address these demands through automated data processing using standardized workflows that produce faster scientific insights. The APS Data Management System coordinates various data related tasks to manage storage, data transfer, metadata cataloging, data processing, and interfaces with tools provided by Globus. Through integration with the Argonne Leadership Computing Facility (ALCF), APS users can efficiently access high-performance computing resources. Standardized workflows have led to reduced computational burdens on scientists and greater accessibility of high performance computing resources. We demonstrate how standardization and collaboration enable scientists to rapidly convert raw data into meaningful scientific results, establishing a streamlined path from data collection to analysis and ultimately to publication.

Parraga, Hannah↗

Enabling Low-Overhead HT-HPC Workflows at Extreme Scale using GNU Parallel

GNU Parallel is a versatile and powerful tool for process parallelization widely used in scientific computing. This paper demonstrates its effective application in high-performance computing (HPC) environments, particularly focusing on its scalability and efficiency in executing large-scale high-throughput high-performance computing (HT-HPC) workflows. Through real-world examples, we highlight GNU Parallel’s performance across various HPC workloads, including GPU computing, container-based workloads, and node-local NVMe storage. Our results on two leading supercomputers, OLCF’s Frontier and NERSC’s Perlmutter, showcase GNU Parallel’s rapid process dispatching ability and its capacity to maintain low overhead even at extreme scales. We explore GNU Parallel’s application in massive parallel file transfers using a scheduled Data Transfer Node (DTN) cluster, emphasizing its broad utility in diverse scientific workflows. Beyond its direct application as a viable workflow manager, GNU Parallel can be employed in conjunction with other workflow systems as a "last-mile" parallelizing driver and as a quick prototyping tool to design and extract parallel profiles from application executions. We then argue that the potential for GNU Parallel to transform workflow management at extreme scales is substantial, paving the way for more efficient and effective scientific discoveries.

Maheshwari, Ketan↗

PIPES (Pipeline for Integrated Projects in Energy Systems) [SWR-24-89]

The Pipeline for Integrated Projects in Energy Systems (PIPES) is a comprehensive project, data, and workflow management tool designed for integrated modeling teams. PIPES facilitates the management of data requirements, tasks, and progress tracking, serving as a higher-level integration layer that works across various data and modeling software. This tool integrates models, data, and tools to perform large-scale, integrated analysis work at scale. PIPES is designed to streamline integrated modeling projects, enhance collaboration, and ensure the quality and efficiency of data management and workflow processes. https://github.com/nrel-pipes/pipes-api https://github.com/nrel-pipes/pipes-web https://github.com/nrel-pipes/nrel-pipes

Gu, Jianli↗