Engineering Papers⌕ Search

Engineering topics

Allcock, William

Publications and source records attributed to Allcock, William.

2023 Operational Assessment Report: Argonne Leadership Computing Facility

In 2004, the U.S. Department of Energy’s (DOE’s) Advanced Scientific Computing Research (ASCR) program established the Leadership Computing Facility (LCF) with a mission to provide the world’s most advanced computational resources to the open science community. The LCF is a huge investment in the nation’s scientific and technological future, inspired by a growing demand for large-scale computing and its impact on science and engineering. The LCF operates two world-class centers in support of open science at Argonne National Laboratory (Argonne) and at Oak Ridge National Laboratory (Oak Ridge) and deploys diverse machines that are among the most powerful systems in the world today. The LCF ranks among the top U.S. scientific facilities delivering impactful science. The work performed at these centers informs policy decisions and advances innovations in far-reaching areas such as energy assurance, ecological sustainability, and global security. The leadership-class systems at Argonne and Oak Ridge operate around the clock every day of the year. The high level of services these centers provide and the exceptional science they produce justify their existence to the DOE Office of Science and the U.S. Congress

97 MATHEMATICS AND COMPUTING↗

Empowering Scientific Discovery Through Computing at the Advanced Photon Source

This paper explores the challenges and solutions for managing and processing the vast amount of data generated by the Advanced Photon Source (APS), a synchrotron light source facility producing ultra-bright x-rays for diverse scientific domains. With 68 experimental beamlines covering materials research, biology, and more, the APS serves a wide user base across academia, government, and industry. The ongoing upgrade of the APS storage ring and installation of new instruments will amplify data generation and processing demands. This paper discusses the approach to address these demands through automated data processing using standardized workflows that produce faster scientific insights. The APS Data Management System coordinates various data related tasks to manage storage, data transfer, metadata cataloging, data processing, and interfaces with tools provided by Globus. Through integration with the Argonne Leadership Computing Facility (ALCF), APS users can efficiently access high-performance computing resources. Standardized workflows have led to reduced computational burdens on scientists and greater accessibility of high performance computing resources. We demonstrate how standardization and collaboration enable scientists to rapidly convert raw data into meaningful scientific results, establishing a streamlined path from data collection to analysis and ultimately to publication.

Parraga, Hannah↗

DRAS: Deep Reinforcement Learning for Cluster Scheduling in High Performance Computing

Cluster schedulers are crucial in high-performance computing (HPC). They determine when and which user jobs should be allocated to available system resources. Existing cluster scheduling heuristics are developed by human experts based on their experience with specific HPC systems and workloads. However, the increasing complexity of computing systems and the highly dynamic nature of application workloads have placed tremendous burden on manually designed and tuned scheduling heuristics. More aggressive optimization and automation are needed for cluster scheduling in HPC. In this work, we present an automated HPC scheduling agent named DRAS (Deep Reinforcement Agent for Scheduling) by leveraging deep reinforcement learning. DRAS is built on a hierarchical neural network incorporating special HPC scheduling features such as resource reservation and backfilling. An efficient training strategy is presented to enable DRAS to rapidly learn the target environment. Once being provided a specific scheduling objective given by the system manager, DRAS automatically learns to improve its policy through interaction with the scheduling environment and dynamically adjusts its policy as workload changes. We implement DRAS into a HPC scheduling platform called CQGym. CQGym provides a common platform allowing users to flexibly evaluate DRAS and other scheduling methods such as heuristic and optimization methods. Furthermore, the experiments using CQGym with different production workloads demonstrate that DRAS outperforms the existing heuristic and optimization approaches by up to 50%.

97 MATHEMATICS AND COMPUTING↗