Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “cluster scheduling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Oak Ridge Computing Academy: An HPC cluster deployment and management pilot

The High Performance Computing Technologies (HPCT) course is a hands-on High Performance Computing (HPC) cluster deployment and management training program offered as part of the International School for Advanced Studies (SISSA) and the International Center for Theoretical Physics (ICTP) Master in High Performance Computing (MHPC) specialization. Here, this training program introduces students to key concepts in cluster configuration. which include networking, software stack provisioning, job scheduling, and monitoring. The publicly available course materials feature several examples and underlying methods that are broadly applicable to cluster deployment and management. This paper discusses the design of a new workforce development program at the Oak Ridge National Laboratory that is based on HPCT, the Oak Ridge Computing Academy (ORCA). The ORCA pilot program was hosted by the Oak Ridge Leadership Computing Facility (OLCF) in Summer 2025. As a part of this discussion, HPCT and ORCA course contents and infrastructure are outlined, ORCA participant experiences are detailed, and potential opportunities for improvement are discussed.

Education↗

Preventive Power Outage Estimation Based on A Novel Scenario Clustering Strategy: Preprint

The increasing occurrence of extreme weather events is challenging the power grid operation. In front of the extreme weather, the system operator is responsible for estimating the power outage and scheduling the restoration resources. This paper proposes an outage evaluation framework to identify the possible unserved load profiles, vulnerable areas, and mobile energy adequacy. The predicted vulnerable lines of an outage prediction model tool are utilized to generate numerous faulted line scenarios. Next, each scenario's nodal unserved load profile is obtained by solving a three-phase restoration model that considers the schedule of repair crews and mobile energy resources. Then, a novel scenario clustering strategy is developed to cluster the unserved load profiles into multiple representative ones for straightforward analysis. Finally, case studies on a distribution system evaluate the damage level brought by extreme weather and verify the effectiveness of the proposed scenario clustering strategy.

mobile energy resources↗

A Hands-On Curriculum for Training in HPC Cluster Deployment and Management

This paper presents the design, methodology, and outcomes of the High-Performance Computing Technologies (HPCT) course, a hands-on training program focused on the system-side of HPC cluster deployment and administration. Delivered as part of the Master in High Performance Computing (MHPC) program, the course introduces students to key concepts in cluster configuration, including networking, software stack provisioning, job scheduling, and monitoring. Initially taught in person, the course was transitioned to an online format during the COVID-19 pandemic. This shift led to the development of openly available instructional material and a flipped-classroom approach that continues to support both in-person and hybrid delivery. All course materials are publicly available at www.hpc.temple.edu/mhpc/hpc-technology/index.html. By documenting the structure, infrastructure, and evolution of HPCT, this paper offers a model for accessible HPC system training that supports workforce development in computational science.

Posada Correa, Fernando [ORNL] (ORCID:000000022565↗

Shaping the FutureWorkforce: Challenges and Lessons Learned in HPC Education from National Labs and Computing Centers

Workforce training at national laboratories and computing centers is essential and typically falls into two categories: foundational training for newcomers and advanced training for experienced users. Foundational topics—such as version control, build systems, and basic HPC usage—are largely transferable across institutions, while cluster-specific training varies due to differences in hardware, job schedulers, and local workflows. Training on emerging technologies is split between hardware-specific content and broadly applicable programming paradigms. Here, to reduce redundancy and increase impact, national labs, computing centers, and vendors are collaborating through initiatives like the HPC Training Working Group to share best practices, co-develop materials, and broaden outreach. These coordinated efforts aim to make HPC training more accessible, scalable, and consistent across the community.

HPC↗

Enhancing Active Distribution Systems Resilience by Fully Distributed Self-Healing Strategy

Distributed restoration can exploit smart grid technologies to enhance the resilience of active distribution networks toward a self-healing smart grid. However, the large number of decision variables, especially the binary ones for reconfiguration, bring challenges to developing scalable distributed distribution service restoration (DDSR) strategies. This paper proposes a fully distributed solution procedure based on the alternating direction method of multipliers (ADMM) for mixed-integer programming problems and applies to develop the DDSR framework. The method consists of relax-drive-polish phases, 1) relaxing binary variables, and applying the convex ADMM as a warm start; 2) driving the solutions toward Boolean values through a proximal operator; 3) fixing the obtained binding binary variables and solving the rest of the problem to polish results and achieve a high-quality suboptimal solution. Then, an autonomous clustering strategy and consensus ADMM are integrated with the proposed method to realize the fully distributed cluster-based framework of DDSR. This framework can first determine DER scheduling and switch status for reconfiguration to energize the out-of-service areas from local faults, and then provide the load restoration solution in a distributed manner for total blackouts in large-scale distribution networks. Furthermore, the effectiveness and scalability of the proposed DDSR framework are demonstrated through testing on the IEEE 123-node, IEEE 8500-node, and synthetic 100k-node test feeders.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Informed Feature Selection for Data Clustering of CSP Plant Production

To make concentrating solar power (CSP) more cost competitive, rigourous optimizations must be run to improve plant design and operations. However, these optimizaitons rely on time consuming annual simulations that solve an electricity dispatch scheduling problem to maximize plant revenue. To reduce the runtime of annual dispatch simulations of CSP plants, a data clustering approach is utilized. This approach assumes that like days of revenue and electricity generation can be identified using weather and price data. Although weather and price are important factors for electricity production, this work investigates how thermal energy storage (TES) inventory at the beginning of a day, denoted as Si, can be used as a supplemental feature to group like days. A framework for creating and training a deep neural network to predict Si is proposed. This model is validated and assessed using eleven sets of testing data that were not used during training. Then, the data clustering approach is performed three seperate times with features of weather and price along with either Si from the neural network, Si from the full annual simulation, or no Si. Ultimately, the results suggest that using Si as an additional clustering feature improves the data clustering simulation accuracy by 1.4%.

Tuman, Matthew J. (ORCID:000900038772051X)↗

SchedInspector: A Batch Job Scheduling Inspector Using Reinforcement Learning

Improving the performance of job executions is an important goal of HPC batch job schedulers, such as minimizing job waiting time, slowdown, or completion time. Such a goal is often accomplished using carefully designed heuristics based on job features, such as job size and job duration. However, these heuristics overlook important runtime factors (e.g., cluster availability and waiting job patterns), which may vary across time and make a previously sound scheduling decision not hold any longer. In this study, we propose a new approach to incorporate runtime factors into batch job scheduling for better job execution performance. The key idea is to add a scheduling inspector on top of the base job scheduler to scrutinize its scheduling decisions. The inspector will take the runtime factors into consideration and accordingly determine the fitness of the scheduled job. It then either accepts the scheduled job or rejects it and asks the base schedulers to try again later. We realize such an inspector, namely SchedInspector, by leveraging the intelligence of reinforcement learning. Through extensive experiments, we show SchedInspector can intelligently integrate the runtime factors into various batch job scheduling policies, including the state-of-the-art one, to gain better job execution performance, such as smaller average bounded job slowdown (up to 69% better) or average job waiting time (up to 52% better), across various real-world workloads. We also show that although rejecting scheduling decisions may leave the resources idle hence affect the system utilization, SchedInspector is able to achieve the job execution performance improvement with marginal impact on the system utilization (typically less than 1%). We consider one key advantage of SchedInspector is it automatically learns to work with and improve existing job scheduling policies without changing them, which makes it promising to serve as a generic enhancer for various batch job scheduling policies.

Zhang, Di↗

Preventive Power Outage Estimation Based on a Novel Scenario Clustering Strategy

The increasing occurrence of extreme weather events is challenging power grid operation. For extreme weather events, the system operator is responsible for estimating the power outages and scheduling the restoration resources. This paper proposes an outage evaluation framework to identify the possible unserved load profiles, vulnerable areas, and mobile energy adequacy. The outputs of an outage prediction model tool are used to generate numerous faulted line scenarios. Next, each scenario's nodal unserved load profile is obtained by solving a three-phase restoration model that considers repair crews and mobile energy resources (MERs). Then, a novel scenario clustering strategy is developed to cluster the unserved load profiles into multiple representative profiles which the system operator can focus on. Finally, case studies on a distribution system evaluate the damage caused by an extreme weather event and verify the effectiveness of the proposed scenario clustering strategy.

MATHEMATICS AND COMPUTING,POWER TRANSMISSION AND D↗

Maskman

SAND2025-04369O Maskman is a user-friendly tool designed to create hex masks, which are essential for optimizing application performance in high-performance computing environments. By converting a list of integers into binary and then hex masks, Maskman simplifies the process of setting application affinity. This ensures that software runs efficiently on specific nodes within a computing cluster. Ideal for researchers and developers, Maskman streamlines the preparation of inputs for HPC schedulers, enhancing resource management and improving overall system performance. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Pase, Douglas [Sandia National Lab. (SNL-CA), Live↗

VC3: Virtual Clusters for Community Computation

A traditional HPC computing facility provides a large amount of computing power but has a fixed environment designed to satisfy local needs. This makes it very challenging for users to deploy complex applications that span multiple sites and require specific application software, scheduling middleware, or sharing policies. This project addressed many of these challenges by making it possible for researchers to easily aggregate and share resources, install custom software environments, and deploy clustering frameworks across multiple HPC facilities through the concept of “virtual clusters”. We designed and implemented a prototype virtual cluster facility that enabled unprivileged users to create dynamic aggregations of computing power across multiple sites, deployed with custom middleware and complex software dependencies. This service is hosted at the University of Chicago and available through the site virtualclusters.org.

97 MATHEMATICS AND COMPUTING↗

OctoFAS: A Two-Level Fair Scheduler That Increases Fairness in Network-Based Key-Value Storage

We identified a fairness problem in a network-based key-value storage system using Intel Storage Performance Development Kit (SPDK) in a multitenant environment. In such an environment, each tenant’s I/O service rate is not fairly guaranteed compared to that of other tenants. To address the fairness problem, we propose OctoFAS, a two-level fair scheduler designed to improve overall throughput and fairness among tenants. The two-level scheduler of OctoFAS consists of (i) inter-core scheduling and (ii) intra-core scheduling. Through inter-core scheduling, OctoFAS addresses the load imbalance problem that is inherent in SPDK on the storage server by dynamically migrating I/O requests from overloaded cores to underloaded cores, thereby increasing overall throughput. Intra-core scheduling prioritizes handling requests from starving tenants over well-fed tenants within core-specific event queues to ensure fair I/O services among multiple tenants. OctoFAS is deployed on a Linux cluster with SPDK. Through extensive evaluations, we found that OctoFAS ensures that the total system throughput remains high and balanced, while enhancing fairness by approximately 10% compared to the baseline, when both scheduling levels operate in a hybrid fashion.

97 MATHEMATICS AND COMPUTING↗

VC3: Virtual Clusters for Community Computation (Final Technical Report)

A traditional HPC computing facility provides a large amount of computing power but has a fixed environment designed to satisfy local needs. This makes it very challenging for users to deploy complex applications that span multiple sites and require specific application software, scheduling middleware, or sharing policies. This project addressed many of these challenges by making it possible for researchers to easily aggregate and share resources, install custom software environments, and deploy clustering frameworks across multiple HPC facilities through the concept of “virtual clusters”. We designed and implemented a prototype virtual cluster facility that enabled unprivileged users to create dynamic aggregations of computing power across multiple sites, deployed with custom middleware and complex software dependencies.

97 MATHEMATICS AND COMPUTING↗

Using containers to speed up development, to run integration tests and to teach about distributed systems

GlideinWMS is a workload manager provisioning resources for many experiments including CMS and DUNE. The software is distributed both as native packages and specialized production containers. Following an approach used in other communities like web development we built our workspaces, system-like containers to ease development and testing. Developers can change the source tree or check out a different branch and quickly reconfigure the services to see the effect of their changes. In this paper, we’ll talk about what differentiates workspaces from other containers. We’ll describe our base system composed of three containers. A one-node cluster including a compute element and a batch system. A GlideinWMS Factory controlling pilot jobs. And a scheduler and Frontend, to submit jobs and provision resources. Additional containers can be used for optional components. This system can easily run on a laptop and we’ll share our evaluation of different container runtimes, with an eye for ease of use and performance. Finally, we’ll talk about our experience as developers and with students. The GlideinWMS workspaces are easily integrated with IDEs like VS Code, simplifying debugging and allowing development and testing of the system also when offline. They simplified the training and onboarding of new team members and Summer interns. And they were useful in workshops where students could have first-hand experience with the mechanisms and components that, in production, run millions of jobs.

Mambelli, Marco↗

Using Containers to Speed Up Development, to Run Integration Tests and to Teach About Distributed Systems

GlideinWMS is a workload manager provisioning resources for many experiments, including CMS and DUNE. The software is distributed both as native packages and specialized production containers. Following an approach used in other communities like web development, we built our workspaces, system-like containers to ease development and testing. Developers can change the source tree or check out a different branch and quickly reconfigure the services to see the effect of their changes. In this paper, we will talk about what differentiates workspaces from other containers. We will describe our base system, composed of three containers: a one-node cluster including a compute element and a batch system, a GlideinWMS Factory controlling pilot jobs, and a scheduler and Frontend to submit jobs and provision resources. Additional containers can be used for optional components. This system can easily run on a laptop, and we will share our evaluation of different container runtimes, with an eye for ease of use and performance. Finally, we will talk about our experience as developers and with students. The GlideinWMS workspaces are easily integrated with IDEs like VS Code, simplifying debugging and allowing development and testing of the system even when offline. They simplified the training and onboarding of new team members and summer interns. And they were useful in workshops where students could have first-hand experience with the mechanisms and components that, in production, run millions of jobs.

Mambelli, Marco [Fermilab] (ORCID:0000000294892681↗

The Empirical Effect of Fleet Optimization on Synchronization and Rebound Effects in Heat Pump Water Heaters

Demand response is a growing concept in light of the internet of things and an increasing need for grid flexibility. Water heaters are one of the preferred devices for providing demand response for grid services and peak management due to their capability to store energy. The efficient use of water heaters for demand response requires consideration of the associated load effects such as synchronization of device schedules and rebound effect. These effects present a significant challenge. Despite the importance of the mentioned effects for water heater queuing and scheduling, there has been no effort to quantify and empirically validate their impact. This study attempts to address this gap by offering two methods - Ward clustering and Euclidean K-means - to evaluate the extent of synchronization in a fleet of 42 water heaters in Atlanta, GA. Using the aforementioned methods on the measured data, we find evidence of convergence of water heater loads as a result of optimization compared to an idle period and analyzed their impact.

demand response↗

Optimizing Performance on Trinity Utilizing Machine Learning, Proxy Applications and Scheduling Priorities

The sheer number of nodes continues to increase in today’s supercomputers, the first half of Trinity alone contains more than 9400 compute nodes. Since the speed of today’s clusters are limited by the slowest nodes, it more important than ever to identify slow nodes, improve their performance if it can be done, and assure minimal usage of slower nodes during performance critical runs. This is an ongoing maintenance task that occurs on a regular basis and, therefore, it is important to minimize the impact upon its users by assessing and addressing slow performing nodes and mitigating their consequences while minimizing down time. These issues can be solved, in large part, through a systematic application of fast running hardware assessment tests, the application of Machine Learning, and making use of performance data to increase efficiency of large clusters. Proxy applications utilizing both MPI and OpenMP were developed to produce data as a substitute for long runtime applications to evaluate node performance. Machine learning is applied to identify underperforming nodes, and policies are being discussed to both minimize the impact of underperforming nodes and increase the efficiency of the system. In this paper, I will describe the process used to produce quickly performing proxy tests, consider various methods to isolate the outliers, and produce ordered lists for use in scheduling to accomplish this task.

97 MATHEMATICS AND COMPUTING↗

Renewable hydrogen and ammonia for combined heat and power systems in remote locations: Optimal design and scheduling

Abstract Using hydrogen (H ) and ammonia (NH ) for renewable energy storage has the potential to enable economical power and heat supply with high renewable penetrations, especially in remote locations which are characterized by high energy costs. In this work we assess the economic competitiveness of renewable combined heat and power (CHP) systems in Mahaka HI, Nantucket MA, and Northwest Arctic Borough (NWAB) AK by optimally designing these systems for scenarios in which power and heat can be purchased over a range of historical energy prices as well as when 100% renewable supply is required. We use a combined optimal design and scheduling model which minimizes annualized net present cost by determining optimal technology selection and size simultaneously with optimal schedules for each period of a system operating horizon aggregated from full year hourly resolution data via a consecutive temporal clustering algorithm. We find that renewable generation meets at least 85% of power demands and 75% of heat demands under the lowest energy prices investigated. Higher conventional energy prices lead to increased renewable penetration which is facilitated by renewable NH as a seasonal energy storage medium, as are 100% renewable CHP systems. NH is used for power generation with heat cogeneration in all three locations, as well as directly for heating in NWAB. On an annual cost basis, NH ‐enabled 100% renewable CHP is only 3% more expensive in Mahaka and NWAB than systems which can purchase energy at the lowest prices, while it is 15% more expensive in Nantucket.

Palys, Matthew J.↗

Condition-Based Maintenance of a Circulating Water System of a Canadian Nuclear Power Plant using Machine Learning and Statistical Tools

Canada Deuterium Uranium pressurized-heavy-water reactors (PHWR) are a type of nuclear power plant that generate clean and reliable energy. The scope of this work is to automate data analysis methodologies to inform a condition-based maintenance strategy of a circulating water system (CWS) of a PHWR. The multiunit CWS provides a continuous supply of water to cool steam condensers, even during transient scenarios, thereby improving the thermal efficiency. This work aims to develop a machine learning (ML) based approach to detect anomalies in heterogeneous data of a CWS in a PHWR to help inform a predictive maintenance strategy. The heterogeneous data include textual and numeric time series data for a PHWR. Natural-language-processing (NLP)-based models are used to analyze textual data contained in work orders and operator logs and an event-timeseries correlation detection method is applied to assist anomalies diagnoses for CWS. An ML model Robust Linear Model (RLM) is also used to remove the seasonal variations in the system variable distributions based on distributions of environmental variables. A machine learning model, Density-Based Spatial Clustering of Applications with Noise (DBSCAN), trained on both original data and data without any seasonal variations will then be used to detect if an anomaly exists. Thus, by moving to an automated methodology to detect, classify, and forecast anomalies, the maintenance strategy would be based on component condition instead of a time-based schedule.

97 - MATHEMATICS AND COMPUTING↗