Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “schedules”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Smart thermostat data-driven U.S. residential occupancy schedules and development of a U.S. residential occupancy schedule simulator

Occupancy schedule is one of the key inputs in Building Energy Modeling (BEM) to reflect the interaction between buildings and occupants. Over the past decades, standardized occupancy schedules, developed mainly by engineering rule-of-thumb, have been widely used in BEM due to its simplicity and lack of real measured occupancy data. However, the BEM community has recognized their association with uncertainty and reliability in simulation results from BEM. This study introduces representative occupancy schedules in the U.S. residential buildings, derived from a large smart thermostat dataset and time-series K-means clustering, and an open-source tool to generate a stochastic residential occupancy schedule. Over 90,000 residential occupancy schedules were estimated from the ecobee Donate Your Data dataset. Then, the representative occupancy schedules were identified through clustering. This study further investigated the impacts of three parameters (day, house type, and state) on residential occupancy schedules. Then, a tool, the Residential Occupancy Schedule Simulator (ROSS), is developed using the representative occupancy schedules derived in this study. Details of this tool are presented in this paper. In conclusion, the derived representative occupancy schedules and the ROSS tool can help improve the energy modeling of residential buildings.

42 ENGINEERING↗

Low Latency and High Data Rate (LLHD) Scheduler: A Multipath TCP Scheduler for Dynamic and Heterogeneous Networks

The scheduler is a crucial component of the multipath transmission control protocol (MPTCP) that dictates the path that a data packet takes. Schedulers are in charge of delivering data packets in the right order to prevent delays caused by head-of-line blocking. The modern Internet is a complicated network whose characteristics change in real-time. MPTCP schedulers are supposed to understand the real-time properties of the underlying network, such as latency, path loss, and capacity, in order to make appropriate scheduling decisions. However, the present scheduler does not take into account all of these characteristics together, resulting in lower performance. We present the low latency and high data rate (LLHD) scheduler, which successfully makes scheduling decisions based on real-time information on latency, path loss, and capacity, and achieves around 25% higher throughput and 45% lower data transmission delay than Linux’s default MPTCP scheduler.

97 MATHEMATICS AND COMPUTING↗

Outage Forecast-Based Preventative Scheduling Model for Distribution System Resilience Enhancement

Distribution system resilience enhancement is an important topic to ensure customers have access to power supply during extreme events. In fact, certain weather-related extreme events can be predicted ahead of time. Therefore, it is important to investigate how to predict grid outages using extreme weather forecasts, and how outage predictions can be incorporated into distribution system resilience enhancement. In this paper, a preventative scheduling model for distribution systems is proposed. The model targets at allocating resources, especially mobile responsive resources such as mobile backup generators and mobile energy storage systems, to prepare for an extreme event in the day-ahead context. To achieve efficient resource allocation and scheduling, a machine learning-based outage prediction module is developed to predict vulnerable or risky segments of the distribution system based on historical operating records and extreme weather event forecast. By integrating the outage prediction results into the scheduling model, optimal resource allocation can be derived to help distribution systems prepare for an upcoming event and improve resilience performance. A real distribution feeder in North Carolina, U.S. is used in the case study to validate the proposed approach.

distributed energy resources↗

Tools And Methods to Analyze Plant Outage Schedule and Assist Schedulers in Improving Outage Resilience

Refueling outages of nuclear power plants (NPPs) are considered one of the most critical phases throughout the plant lifetime. In such instances, tens of thousands of activities (e.g., maintenance, surveillance) are performed in a short amount of time (typically 2-3 weeks unless major backfitting or modernization projects are carried out) by a large number of crews (e.g., electricians, mechanics) that are hired as contractors. As a consequence, a plant outage can be expensive not only in terms of costs (e.g., contractor labor, material), but also in terms of loss generation since the plant is taken off the grid during the full outage duration (an indicative metric is about 1.2M$/day of loss of revenue). Thus, there is a continuous need to decrease the economic impact of outages on plant finances. This can be done by: decreasing the frequency of plant outages (e.g., from 18 to 24 months), reducing the time to complete the outage, and reducing the risk of outage delays. The Optimization of Outage Activities project under the Risk Informed Systems Analysis Pathway (RISA) sponsored by Department of Energy (DOE) Light Water Reactor Sustainability (LWRS) Program focuses on developing tools and methods to support NPPs with outage schedule optimization. The developed tools and methods are designed to analyze plant outage schedule with the goal of identify critical elements in the schedule that might pose a high risk of delays. These methods and tools can be considered resource-centric in the sense that they address outage challenges as a resource optimization problem. In this context, resources are either time and crews; outage delays occurs when either (or both) resources are insufficient to complete the set of tasks assigned at a specific time instant of the outage. This report provides details on how plant resources (time and crews) can be allocated in such a way that delays are minimized. In this respect, two classes of methods have been developed: the first one focuses on the time resource and how variability of the time to complete outage tasks may impact outage delays. The second one integrates available resources to assess when dailies activities should be performed such that the risk of outage delays are minimized.

97 MATHEMATICS AND COMPUTING↗

SchedInspector: A Batch Job Scheduling Inspector Using Reinforcement Learning

Improving the performance of job executions is an important goal of HPC batch job schedulers, such as minimizing job waiting time, slowdown, or completion time. Such a goal is often accomplished using carefully designed heuristics based on job features, such as job size and job duration. However, these heuristics overlook important runtime factors (e.g., cluster availability and waiting job patterns), which may vary across time and make a previously sound scheduling decision not hold any longer. In this study, we propose a new approach to incorporate runtime factors into batch job scheduling for better job execution performance. The key idea is to add a scheduling inspector on top of the base job scheduler to scrutinize its scheduling decisions. The inspector will take the runtime factors into consideration and accordingly determine the fitness of the scheduled job. It then either accepts the scheduled job or rejects it and asks the base schedulers to try again later. We realize such an inspector, namely SchedInspector, by leveraging the intelligence of reinforcement learning. Through extensive experiments, we show SchedInspector can intelligently integrate the runtime factors into various batch job scheduling policies, including the state-of-the-art one, to gain better job execution performance, such as smaller average bounded job slowdown (up to 69% better) or average job waiting time (up to 52% better), across various real-world workloads. We also show that although rejecting scheduling decisions may leave the resources idle hence affect the system utilization, SchedInspector is able to achieve the job execution performance improvement with marginal impact on the system utilization (typically less than 1%). We consider one key advantage of SchedInspector is it automatically learns to work with and improve existing job scheduling policies without changing them, which makes it promising to serve as a generic enhancer for various batch job scheduling policies.

Zhang, Di↗

Principled Schedulability Analysis for Distributed Storage Systems Using Thread Architecture Models

In this article, we present an approach to systematically examine the schedulability of distributed storage systems, identify their scheduling problems, and enable effective scheduling in these systems. We use Thread Architecture Models (TAMs) to describe the behavior and interactions of different threads in a system, and show both how to construct TAMs for existing systems and utilize TAMs to identify critical scheduling problems. We specify three schedulability conditions that a schedulable TAM should satisfy: completeness, local enforceability, and independence; meeting these conditions enables a system to easily support different scheduling policies. We identify five common problems that prevent a system from satisfying the schedulability conditions, and show that these problems arise in existing systems such as HBase, Cassandra, MongoDB, and Riak, making it difficult or impossible to realize various scheduling disciplines. We demonstrate how to address these schedulability problems using both direct and indirect solutions, with different trade-offs. To show how to apply our approach to enable scheduling in realistic systems, we develop Tamed-HBase and Muzzled-HBase, sets of modifications to HBase that can realize the desired scheduling disciplines, including fairness and priority scheduling, even when presented with challenging workloads.

Computer Science↗

Decentralized Distributed Proximal Policy Optimization (DD-PPO) for High Performance Computing Scheduling on Multi-User Systems

Resource allocation in High Performance Computing (HPC) environments presents a complex and multifaceted challenge for job scheduling algorithms. Beyond the efficient allocation of system resources, schedulers must account for and optimize multiple performance metrics, including job wait time and system throughput. Traditional heuristic-based scheduling algorithms increasingly struggle and lack the efficiency needed to meet the demands and address the complexity and scale of modern HPC systems. Consequently, recent research efforts have focused on leveraging advancements in Artificial Intelligence (AI) and Deep Learning (DL), particularly Reinforcement Learning (RL), to develop more adaptable and intelligent scheduling strategies. Previous RL-based scheduling approaches have explored a range of algorithms, from Deep Q-Networks (DQN) to Proximal Policy Optimization (PPO), and more recently, hybrid methods that integrate Graph Neural Networks (GNNs) with RL techniques. However, a common limitation across these methods is their reliance on relatively small datasets, with few methods being evaluated using large-scale, multi-million-job trace datasets representative of real-world HPC workloads. Moreover, existing RL schedulers face scalability issues due to centralized policy updates, which hinder training efficiency and performance when applied to large datasets. This study introduces a novel RL-based scheduler utilizing Decentralized Distributed Proximal Policy Optimization (DD-PPO) algorithm, which supports large-scale distributed training across multiple workers without requiring parameter synchronization at every step. By eliminating reliance on centralized updates to a shared policy, the DD-PPO scheduler enhances scalability, training efficiency, and sample utilization. Experimental validation using a large real-world dataset containing over 11.5 million job traces collected from petascale HPC systems over six years assesses the influence of dataset scale on training effectiveness and compares DD-PPO performance to traditional and advanced scheduling approaches. The experimental results demonstrate improved scheduling performance in comparison to both heuristic-based schedulers and existing RL-based scheduling algorithms.

AI↗

DRAS: Deep Reinforcement Learning for Cluster Scheduling in High Performance Computing

Cluster schedulers are crucial in high-performance computing (HPC). They determine when and which user jobs should be allocated to available system resources. Existing cluster scheduling heuristics are developed by human experts based on their experience with specific HPC systems and workloads. However, the increasing complexity of computing systems and the highly dynamic nature of application workloads have placed tremendous burden on manually designed and tuned scheduling heuristics. More aggressive optimization and automation are needed for cluster scheduling in HPC. In this work, we present an automated HPC scheduling agent named DRAS (Deep Reinforcement Agent for Scheduling) by leveraging deep reinforcement learning. DRAS is built on a hierarchical neural network incorporating special HPC scheduling features such as resource reservation and backfilling. An efficient training strategy is presented to enable DRAS to rapidly learn the target environment. Once being provided a specific scheduling objective given by the system manager, DRAS automatically learns to improve its policy through interaction with the scheduling environment and dynamically adjusts its policy as workload changes. We implement DRAS into a HPC scheduling platform called CQGym. CQGym provides a common platform allowing users to flexibly evaluate DRAS and other scheduling methods such as heuristic and optimization methods. Furthermore, the experiments using CQGym with different production workloads demonstrate that DRAS outperforms the existing heuristic and optimization approaches by up to 50%.

97 MATHEMATICS AND COMPUTING↗

Simulation-based assessment on stochastic load scheduling for building cooling systems

Here, to fill knowledge gaps related to stochastic load scheduling, we performed a comprehensive evaluation of the stochastic load scheduling for building cooling systems. Specifically, we studied the common uncertain variables in the load scheduling process for building cooling systems and categorized those variables based on their dynamic patterns. We then developed a generic stochastic load scheduling framework and applied it to building cooling systems that served a simulated community. This community consists of 100 heterogeneous houses and serves as a virtual testbed for evaluating the performance of stochastic load scheduling. In this evaluation, we considered representatives of uncertain variables with different dynamic patterns and included 100 realizations of the considered uncertainty in the evaluation to better catch the probability distribution of the control performance. The evaluation results suggest that deterministic load scheduling can reduce the operating energy cost by 18% but its performance can be affected by uncertainty. Stochastic load scheduling can further decrease the operating energy cost under uncertainty compared to deterministic load scheduling. We also found that the effectiveness of stochastic load scheduling in handling uncertainty is not directly associated with the number of uncertainty scenarios that are considered in its formulation.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Quantum/AI Topology-Aware Latency-Adaptive HPC Workflow Scheduling Optimization

The growing demand for more powerful high-performance computing (HPC) systems has led to a steady rise in energy consumption by supercomputing worldwide. This study is focused on comparing our Application-Topology Mapper (ATMapper) to the popular Simple Linux Utility for Resource Management (SLURM) for the purpose of exploring methods that can further optimize job-scheduling within HPC systems. ATMapper is an Artificial-Intelligence based approach to job-scheduling that is currently being enhanced with quantum annealing (QA) to generate optimal schedules faster. We are applying QA to speedup our ATMapper process to achieve higher computing efficiency, thereby reducing HPC energy consumption. Here, we examine how four job-scheduling approaches perform in processor node assignment when using an example network architecture of 4 interconnected nodes. Using a specialized script, we are assessing the schedule of a computation flow with 11 interdependent tasks. The data movements among nodes were tracked to count for the number of interactions (network hops) between nodes needed to complete the tasks. The total number of hops and the job completion time were then used to quantify the efficiency of the different mapping approaches. In addition to SLURM, we also compare our ATMapper to the QA-enabled LBNL TIGER and the D-Wave Distributed Computing processor assignment approaches. The preliminary results showed that our topology-aware, latency-adaptive ATMapper is significantly more efficient when compared to the other scheduling approaches due to its load-imbalance network allocation. The scheduler displayed a computing efficiency of 53% by performing significantly fewer network hops than its alternatives. By reducing the number of hops, ATMapper was able to perform all 11 tasks by using only 3 nodes out of given 4. This research indicates the potential to use QA/AI for HPC job-scheduling. Later, we will test a SLURM simulator program to draw further comparisons on the effectiveness of ATMapper's scheduling approach. The results of this comparison will serve as a baseline for later improving SLURM's performance using a QA-enhanced ATMapper approach.

Caraveo, Braulio [University of Huston - Clear Lak↗

Occupancy schedule development and its effect on OpenStudio prototype college building model

College buildings have unique characteristics compared with school buildings. Therefore, defining the realistic occupancy schedule in a prototype college building has significant research opportunities. In this study, the actual operating schedules of each space type were collected and generated based on the class reservation schedule and compared with the previous reference schedule (primary/secondary school). Here, the schedules were analyzed for their effect on the OpenStudio prototype college building model. The findings highlight that the use of a typical school building schedule in a college building impairs the granularity of information. The analysis shows significant differences between the previous occupancy schedule and the updated occupancy schedule of the college building, leading to a considerable decrease in occupancy density. Furthermore, the effect of these occupancy pattern changes on the prototype building model is examined. The variations were observed in minimum ventilation requirements, the average mechanical ventilation rate, and energy consumption attributed to changes in occupancy density.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Multi-Timescale Integrated Dynamics and Scheduling for Solar (MIDAS-Solar) (Final Technical Report)

Solar photovoltaic (PV) installations have experienced unprecedented growth in the United States. PV will become not only an energy producer but also a necessary provider of ancillary services at multiple timescales. Conventional methods to simulate power systems operations - such as long-term production simulation (which typically considers schedules from hours to minutes by using an optimization framework) and short-term transient studies (which simulate dynamics from seconds to sub-seconds using state variables and differential equations) - are not sufficient for studying the multiple-timescale variation of solar generation and its impact on system reliability. Long-term system economics and short-term system dynamics are highly coupled, particularly when the penetration level of renewable generation is extremely high, because the uncertainty and variability of solar generation will impact both power system steady-state and dynamic performance. This project helps meet and exceed the U.S. Department of Energy Office of Energy Efficiency and Renewable Energy Solar Energy Technologies Office goal of systems integration by directly addressing this stability and reliability challenge for power grid planning and operation. We have developed a temporally comprehensive, closed-loop simulation model, named Multi-timescale Integrated Dynamics and Scheduling (MIDAS), that seamlessly simulates power system operations from economic scheduling (day-ahead to hours) to dynamic response analysis (seconds to sub-seconds). For schedules with very high levels of inverter-based resources (IBRs), up to and including 100%, the stability of grid controls has been evaluated through electromagnetic transient (EMT) simulations and power-hardware-in-the-loop (PHIL) simulations of key transient events at key schedule points. Specifically, MIDAS provides: 1) a closed-loop simulation framework for simulating timescales from economic scheduling to dynamic stability analysis; 2) machine learning-based stability assessment; 3) EMT modeling and analysis for large-scale power systems; 4) MIDAS PHIL test bed. We worked with Hawaii Electric Companies to apply the MIDAS study framework to a Maui grid study. The entire island's transmission system was modeled in detail - from a yearly scheduling model, to a second-level frequency dynamic model, down to a sub-second-scale EMT model to address critical stability issues. The project demonstrated how MIDAS can help system planners and operators assess system reliability and stability while the power grid is marching toward a high-renewable, high-IBR future. In this Maui grid study, we found that 100% instantaneous IBR operation is achievable in EMT simulation and PHIL testing, and grid planners and operators might need new analysis/simulation tools to assess grid reliability and stability in the scheduling stage. MIDAS will bring Maui and other systems closer to 100% clean and stable energy futures. (In this study, we examined transient stability. Other topics necessary for 100% IBR operation, such as protection and resource adequacy, were not examined.)

100% Renewables↗

US 2023/0182605 A1 Network constraint energy management system for electric vehicle depot charging and scheduling

Network constraint energy management system for electric vehicle (EV) depot charging and scheduling. In an embodiment, a power schedule is received from an economic dispatch application for a charging depot comprising EV charging station(s) and distributed energy resource(s). The power schedule may be simulated on a distribution network model of the charging depot, according to load flow analysis, to determine whether any grid-code violations occur. In response to the detection of violation(s), a constraint may be generated for each violating node, and the economic dispatch application may be re-executed with the constraint(s) to produce a new power schedule, until no violations are detected. When not all load demand can be satisfied by the power schedule, a charging schedule may be adjusted to ensure that critical energy requirements are satisfied. The final power and charging schedules may be used to schedule and control power generation and charging in the charging depot.

Hafiz, Faeza↗

End-Use Savings Shapes Measure Documentation: Dispatch Schedule Generation for Demand Flexibility Measures

This supplemental document describes the methodology used for determining the dispatch timing of various EUSS demand flexibility measures. Demand flexibility measures are designed to reduce/dispatch electricity demand in buildings during especially beneficial/critical times. The method used in this work utilizes predictions of building loads to generate a schedule that reflects the periods when the building's daily peak load occurs to support decision making in demand flexibility measures. The dispatch schedule generation method described in this document creates an hourly schedule that includes a load dispatch (peak) window for each day for a whole year based on load prediction, with options using different prediction methods: perfect prediction, bin-sampling method, fixed schedule, and outdoor air temperature (OAT)-based prediction method. The perfect prediction method performs a simulation to obtain the annual load profile as predicted load, representing the scenario of perfect load prediction. The bin-sampling method (1) categorizes days into representative bins by temperature characteristics, (2) performs simulations on sample days from each of those bins to create representative (or predicted) load, and (3) assigns representative loads for all days in a year based on the bin categorization. The fixed schedule method defines uniform start and end time of peak window with assumed fixed daily peak time, for all days in a season or a year. The OAT-based prediction method uses the statistics of OAT (minimum and maximum) as the indicators of peak load, with specified delay response time from building loads to temperature. Given the load prediction, daily peak periods are determined as a time window with specified length in each day that include the predicted daily peak load and with a secondary rule such as maximizing energy saving potential. The dispatch schedule generation method is not a standalone measure and is intended to be combined with other demand flexibility measures that could leverage the peak schedule and apply demand controls on specific systems or devices for demand response, such as measures described in "Measure Documentation - Thermostat Control for Load Shedding" and "Measure Documentation - Thermostat Control for Load Shifting".

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

OctoFAS: A Two-Level Fair Scheduler That Increases Fairness in Network-Based Key-Value Storage

We identified a fairness problem in a network-based key-value storage system using Intel Storage Performance Development Kit (SPDK) in a multitenant environment. In such an environment, each tenant’s I/O service rate is not fairly guaranteed compared to that of other tenants. To address the fairness problem, we propose OctoFAS, a two-level fair scheduler designed to improve overall throughput and fairness among tenants. The two-level scheduler of OctoFAS consists of (i) inter-core scheduling and (ii) intra-core scheduling. Through inter-core scheduling, OctoFAS addresses the load imbalance problem that is inherent in SPDK on the storage server by dynamically migrating I/O requests from overloaded cores to underloaded cores, thereby increasing overall throughput. Intra-core scheduling prioritizes handling requests from starving tenants over well-fed tenants within core-specific event queues to ensure fair I/O services among multiple tenants. OctoFAS is deployed on a Linux cluster with SPDK. Through extensive evaluations, we found that OctoFAS ensures that the total system throughput remains high and balanced, while enhancing fairness by approximately 10% compared to the baseline, when both scheduling levels operate in a hybrid fashion.

97 MATHEMATICS AND COMPUTING↗

Job Scheduler-Driven Power Gateway for High Performance Computing

Power gateways in the form of a microgrid can incorporate multiple distributed energy resources (DER) in either grid forming or grid following mode and support high performance computing (HPC) power profiles including the large load-follow requirements observed in multi-user HPC systems. The microgrid’s flexibility to operate in either grid forming or grid following mode and to actively switch between these modes enables baseline power from multiple non-baseline DER while maintaining high power quality metrics for the HPC system. But this enormous flexibility in demand response and time of use shifting is generally programmed independently of any integration with an HPC job scheduler which can better inform the load shaping by the microgrid. While there are many existing approaches where the HPC job scheduler takes in information from the grid to make queue scheduling decisions, this work takes the opposite view and explores a scheduler where the jobs in the queue can directly impact the settings of the grid. Several HPC scheduler strategies are tested where the jobs in the queue directly impact the settings of a microgrid designed for HPC operation which is driving a datacenter with three classes of HPC architectures. The scheduler operation is shown using a microgrid with 64 kW of solar capacity and 320 kWh of battery over a period of 21 days operating with significant low-follow swings, a throttled grid, cloudy conditions, switching between grid following and grid forming modes, and a wide range of battery states-of-charge all while maintaining high quality power metrics. The scheduler provides a mechanism for the job queue to directly impact a power gateway like a microgrid and to improve HPC power outcomes such as maximizing renewable energy usage

microgrid↗