Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Scheduling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Optimization of distributed compute resources utilization in the CMS Global Pool

The CMS Submission Infrastructure is the primary system for managing computing resources for CMS workflows, including data processing, simulation, and analysis. It integrates geographically distributed resources from Grid, HPC, and cloud providers into federated pools managed by HTCondor and Glidein- WMS, for a total of around 500k CPU cores. This system dynamically manages workloads based on priorities defined by the collaboration. Additionally, CMS scheduling strategies must be flexible to handle multiple concurrent workloads while considering changing processing demands and resource availability from various providers.Efficient utilization of vast amounts of distributed compute resources is a key element for the success of the scientific programs of the LHC experiments. Optimizing the system is essential to maximize resource efficiency and fully utilize the distributed computing power. The CMS Submission Infrastructure team thus systematically investigates sources of inefficiency in workload scheduling to reduce their impact. In addition, a strategy of pilot overloading has been introduced to compensate for other inefficiency sources, thereby optimizing resource utilization and enhancing computational throughput.

Mascheroni, Marco [UC, San Diego (main)]↗

Guidance for Developing Digital Twins for Online Condition Monitoring of Nuclear Power Plant Components

Online condition monitoring is an area of active research that may enable optimized scheduling, maintenance, and safety of nuclear power plant components, reducing unnecessary derates while simultaneously improving operational capacity. Digital twins (DTs) are one avenue to conduct online condition monitoring and are currently being explored by national laboratories and universities alike. DTs for online condition monitoring are, in essence, state concurrent models that emulate a physical process which predicts a parameter and compares it against a measured value. The promise of DT is that they may provide additional insights by combining and interpreting various sources of information and may be used for preventative maintenance scheduling optimization or early fault detection. DTs for condition monitoring are projected to be valuable for meeting requirements under 10 CFR 50.55a and 10 CFR 50.65. However, DT technologies are still under significant development and the process for developing a DT for condition monitoring has not been formalized. Therefore, in this work, we present an initial framework for developing a DT, discuss and review the various challenges and considerations for DT deployment, and identify the opportunities that a DT can improve. Here, the presented framework is intended to help developers formulate a strategy when approaching DT development for condition monitoring. A DT use case for a reactor coolant pump is presented to demonstrate the proposed framework.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Characterization and Optimization of the Fitting of Quantum Correlation Functions

This case study presents a characterization and optimization of an application code for extracting parton distribution functions from high energy electron-proton scattering data. Profiling this application code reveals that the phase-space density computation accounts for 93% of the overall execution time for a single iteration on a single core. When executing multiple iterations in parallel on a multicore system, the application spends 78% of its overall execution time idling due to load imbalance. We address these issues by first transforming the application code from Python to C++ and then tackling the application load imbalance via a hybrid scheduling strategy that combines dynamic and static scheduling. These techniques result in a 62% reduction in CPU idle time and a 2.46x speedup in overall execution time per node. In addition, the typically enabled power-management mechanisms in supercomputers (e.g., AMD Turbo Core, Intel Turbo Boost, and RAPL) can significantly impact intra-node scalability when more than 50% of the CPU cores are used. This finding underscores the importance of understanding system interactions with power management, as they can adversely impact application performance, and highlights the necessity of intra-node scaling tests to identify performance degradation that inter-node scaling tests might otherwise overlook.

Chuang, Pi-Yueh [Virginia Tech,Dept. of Computer S↗

Urban-Scale Control of School Bus Fleet Charging and Discharging Strategies Using Single and Multi-Stage Optimization

This paper presents a dual-strategy approach to optimizing charging and discharging schedules for school bus fleets, using the limited charging infrastructure effectively. We aim to ensure that each bus is fully charged for daily operations and aids in grid stability during peak demand. The first strategy utilizes linear programming to schedule overnight charging at available station sockets and strategic discharging during peak periods, efficiently coordinating limited resources. The second strategy employs metaheuristic techniques for continuous optimization, focusing on precise power requirements and offering greater flexibility than the linear model.

Selim, Alaa↗

Electrifying Airport GSE: Monte Carlo Grid Impacts

Airports globally are shifting from ICE-powered to electric Ground Support Equipment (eGSE) to enhance efficiency, reduce operational costs, and improve operator health. Leveraging predictable routes, flat terrain, and low operational speeds, airports provide ideal conditions for electrification. This study evaluates freight GSE electrification at Dallas-Fort Worth International Airport (DFW), USA, using the Agile@ platform, which integrates three analytical methods: Freight Facility Model (FFM), Activity-Structure-Intensity-Fuel (ASIF), and Monte Carlo simulations. Results from 10,000 simulations indicate modest but critical increases in electricity demand and significant variability in GSE energy consumption. These insights emphasize the importance of data-driven scheduling, targeted maintenance, and strategic infrastructure planning. For high-uncertainty scenarios, airports are advised to deploy buffer energy storage systems (battery banks), implement demand-response charging strategies, schedule flexible workforce shifts, and prioritize proactive maintenance-particularly for equipment with higher operational uncertainty, such as tug tractors with trailers. Agile@ thus offers a robust, scalable, and data-driven framework to optimize long-term GSE planning and enhance reliability across diverse airport environments.

Bose, Ranjan [ORNL] (ORCID:0009000791026327)↗

Ensemble Simulations on Leadership Computing Systems

Scientific productivity can be enhanced through workflow management tools, relieving large High Performance Computing (HPC) system users from the tedious tasks of scheduling and designing the complex computational execution of scientific applications. This paper presents a study on the usage of ensemble workflow tools to accelerate science using the Summit and Frontier supercomputing systems. The research aims to connect science domain simulations using Oak Ridge Leadership Computing Facility (OLCF) supercomputing platforms with ensemble workflow methods in order to accelerate HPC-enabled discovery and boost scientific impact. We present the coupling, porting and optimization of Radical-Cybertools on three applications: Chroma, NAMD and LAMMPS. The tools augment traditional HPC monolithic runs with a pilot scheduler. Lessons-learned are discussed for physics, biology and materials science applications. We discuss intrinsic limitations of coupling and porting ensemble workflow tools to applications that run on large HPC systems. The origins of technical challenges and their solutions developed during the implementation process are discussed. Data management strategies, OLCF’s policies for ensembles, and natively supported workflow tools are also summarized.

Georgiadou, Antigoni [ORNL] (ORCID:000000020977631↗

Flexible Ramping Product Procurement in Day-Ahead Markets

Flexible ramping products (FRPs) emerge as a promising instrument for addressing steep and uncertain ramping needs through market mechanisms. Initial implementations of FRPs in North American electricity markets, however, revealed several shortcomings in existing FRP designs. Here, in many instances, FRP prices failed to signal the true value of ramping capacity, most notably evident in zero FRP prices observed in a myriad of periods during which the system was in acute need for rampable capacity. These periods were marked by scheduled but undeliverable FRPs, often calling for operator out-of-market actions. On top of that, the methods used for procuring FRPs have been primarily rule-based, lacking explicit economic underpinnings. In this paper, we put forth an alternative framework for FRP procurement, which seeks to set FRP requirements and schedule FRP awards such that the expected system operation cost is minimized. Using real-world data from U.S. ISOs, we showcase the relative merits of the framework in (i) reducing the total system operation cost, (ii) improving price formation, (iii) enhancing the the deliverability of FRP awards, and (iv) reducing the need for out-of-market actions.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Flexibility Options: A Proposed Product for Managing Imbalance Risk

The presence of variable renewable energy resources with uncertain outputs in day-ahead electricity markets results in additional balancing needs in real-time. Addressing those needs cost-effectively and reliably within a competitive market with unbundled products is challenging as both the demand for and the availability of flexibility depends on day-ahead energy schedules. Existing approaches for reserve procurement usually rely either on oversimplified demand curves that do not consider how system conditions that particular day affect the value of flexibility, or on bilateral trading of hedging instruments that are not co-optimized with day-ahead schedules. This article proposes a new product, ‘Flexibility Options', to address these two limitations. The demand for this product is endogenously determined in the day-ahead market and it is met cost-effectively by considering real-time supply curves for product providers, which are co-optimized with the energy supply. As we illustrate with numerical examples and mathematical analysis, the product addresses the hedging needs of participants with imbalances cost-effectively, provides a less intermittent revenue stream for participants with flexible outputs, promotes value-driven pricing of flexibility, and ensures that the system operator is revenue-neutral. This article provides a comprehensive design that can be further tested and applied in large-scale systems.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

IRIS: A Performance-Portable Framework for Cross-Platform Heterogeneous Computing

From edge to exascale, computer architectures are becoming more heterogeneous and complex. The systems typically have fat nodes, with multicore CPUs and multiple hardware accelerators such as GPUs, FPGAs, and DSPs. This complexity is causing a crisis in programming systems and performance portability. Several programming systems are working to address these challenges, but the increasing architectural diversity is forcing software stacks and applications to be specialized for each architecture. As we show, all of these approaches critically depend on their software framework for discovery, execution, scheduling, and data orchestration. To address this challenge, we believe that a more agile and proactive software framework is essential to increase performance portability and improve user productivity. To this end, we have designed and implemented IRIS: a performance-portable framework for cross-platform heterogeneous computing. IRIS can discover available resources, manage multiple diverse programming platforms (e.g., CUDA, Hexagon, HIP, Level Zero, OpenCL, OpenMP) simultaneously in the same execution, respect data dependencies, orchestrate data movement proactively, and provide for user-configurable scheduling. To simplify data movement, IRIS introduces a shared virtual device memory with relaxed consistency among different heterogeneous devices. IRIS also adds an automatic kernel workload partitioning technique using the polyhedral model so that it can resize kernels for a wide range of devices. Our evaluation on three architectures, ranging from Qualcomm Snapdragon to a Summit supercomputer node, shows that IRIS improves portability across a wide range of diverse heterogeneous architectures with negligible overhead.

97 MATHEMATICS AND COMPUTING↗

Hybrid PDES Simulation of HPC Networks Using Zombie Packets

Although high-fidelity network simulations have proven to be reliable and cost-effective tools to peer into architectural questions for high-performance computing (HPC) networks, they incur a high resource cost. The time spent in simulating a single millisecond of network traffic in the highest detail can take hours, even for static, well-behaved traffic patterns such as uniform random. Surrogate models offer a significant reduction in runtime, yet they cannot serve as complete replacements and should only be used when appropriate. Thus, there is a need for hybrid modeling, where high-fidelity simulation and surrogates run side-by-side. Here, we present a surrogate model for HPC networks in which: packets bypass the network, while the network state is left untouched, i.e., suspended. To bypass the network, we use historical data to estimate the arrival time at which every packet should be scheduled at; to suspend the network, all in-flight packets are scheduled to arrive at their destinations, and are kept in the system to awaken as zombies when switching back to high-fidelity. Speedup for a hybrid model is relative to the proportion of surrogate to high-fidelity. This light-weight surrogate obtained up to 76× speedup. Keeping the zombies in the network showed an increase in the accuracy of the high-fidelity simulation on restart when compared to restarting the network from an empty state.

HPC networks↗

Q-IRIS: The Evolution of the IRIS Task-Based Runtime to Enable Classical-Quantum Workflows

Extreme heterogeneity in emerging HPC systems are starting to include quantum accelerators, motivating runtimes that can coordinate between classical and quantum workloads. We present a proof-of-concept hybrid execution framework integrating the IRIS asynchronous task-based runtime with the XACC quantum programming framework via the Quantum Intermediate Representation Execution Engine (QIR-EE). IRIS orchestrates multiple programs written in the quantum intermediate representation (QIR) across heterogeneous backends (including multiple quantum simulators), enabling concurrent execution of classical and quantum tasks. Although not a performance study, we report measurable outcomes through the successful asynchronous scheduling and execution of multiple quantum workloads. To illustrate practical runtime implications, we decompose a four-qubit circuit into smaller subcircuits through a process known as quantum circuit cutting, reducing per-task quantum simulation load and demonstrating how task granularity can improve simulator throughput and reduce queueing behavior -- effects directly relevant to early quantum hardware environments. We conclude by outlining key challenges for scaling hybrid runtimes, including coordinated scheduling, classical-quantum interaction management, and support for diverse backend resources in heterogeneous systems.

Miniskar, Narasinga Rao [ORNL] (ORCID:000000018259↗

EVI-EnSitePy (Electric Vehicle Infrastructure – Energy Estimation and Site Optimization Tool in Python) [EVI-X Modeling Suite] [SWR-25-07]

EVI-EnSitePy is a comprehensive agent-based tool designed for the analysis and design of high-power charging sites, encompassing a wide array of site agents including Electric Vehicles (EVs), chargers, energy storage units (ESS), renewable energy resources (DER), and loads. This versatile tool offers diverse functionalities and a modular modeling approach, allowing detailed configuration of agents based on power ratings, port numbers, energy capacities, demand requirements, charger interfaces, and flexibility to customize the tool for project specific goals. By simulating agent interactions and employing various metrics, EVI-EnSitePy enables the assessment of site performance, exploration of energy management systems (EMS), and implementation of innovative EV charging policies. Utilizing EV charge schedules and arrival states, the tool performs thorough charging site simulations, with outputs consisting of agent and site-level power profiles and statistical metrics. Employing a tree graph structure, EVI-EnSitePy supports nested site structures and power distribution modeling. The tool's ability to generate charging schedules deterministically or via stochastic analysis further enhances its versatility. Through its features and capabilities, EVI-EnSitePy offers a powerful platform for informed decision-making in the realm of high-power charging site design and operation.

Jackson, Derek [National Renewable Energy Laborato↗

EASY-SHIFT v Alpha

The software is a generic, price- and load-responsive control algorithm integrating heat pumps with thermal energy storage. The algorithm leverages simple models of the system and easily accessible data to schedule operation of heat pumps and thermal energy storage in ways that minimize the cost of operating the heating/cooling system. This tool is specifically designed to be easy to interact with, and something that industry partners are able to adopt. There are two current state of the art approaches. Industry tends to develop very simple algorithms, with predetermined schedules that are not capable of changing operation in response to changes in operating environment. For example, a control designed to avoid high-price electricity from 5-8 PM will not be able to adapt if the high-price period changes to 4-9 PM. Academia commonly develops algorithms called Model predictive control (MPC). MPC requires extensive data and highly trained staff to develop a specific type of simulation model of the building, connect the building to optimization algorithms, and leverage powerful computers. Industry, with limited time/finance budgets for any project, is resistant to adopting MPC due to the associated high complexity and cost.

Grant, Peter [Lawrence Berkeley National Laborator↗

Exascale workflow applications and middleware: An ExaWorks retrospective

Exascale computers offer transformative capabilities to combine data-driven and learning-based approaches with traditional simulation applications to accelerate scientific discovery and insight. However, these software combinations and integrations are difficult to achieve due to the challenges of coordinating and deploying heterogeneous software components on diverse and massive platforms. Here, we present the ExaWorks project, which addresses many of these challenges. We developed a workflow Software Development Toolkit (SDK), a curated collection of workflow technologies that can be composed and interoperated through a common interface, engineered following current best practices, and specifically designed to work on HPC platforms. ExaWorks also developed PSI/J, a job management abstraction API, to simplify the construction of portable software components and applications that can be used over various HPC schedulers. The PSI/J API is a minimal interface for submitting and monitoring jobs and their execution state across multiple and commonly used HPC schedulers. We also describe several leading and innovative workflow examples of ExaWorks tools used on DOE leadership platforms. Furthermore, we discuss how our project is working with the workflow community, large computing facilities, and HPC platform vendors to address the requirements of workflows sustainably at the exascale.

97 MATHEMATICS AND COMPUTING↗

Digital-Twin-Enabling Technologies for Online Condition Monitoring of Nuclear Power Plant Components

Online condition monitoring is an area of active research that may enable optimized scheduling, maintenance, and safety of nuclear power plant components, reducing unnecessary derates while simultaneously improving operational capacity. Digital twins (DTs) are one avenue to conduct online condition monitoring and are currently being explored by national laboratories and universities alike. DTs for online condition monitoring are, in essence, state concurrent models that emulate a physical process which predicts a parameter and compares it against a measured value. A DT’s goal is to provide additional insights by combining and interpreting various sources of information for preventative maintenance scheduling optimization or early fault detection. DTs for condition monitoring are projected to be valuable for meeting requirements under 10 CFR 50.55a, “Codes and Standards,” and 10 CFR 50.65, “Requirements for Monitoring the Effectiveness of Maintenance at Nuclear Power Plants”. However, DT technologies are still under significant development, and the process for developing a DT for condition monitoring has not been formalized. Therefore, in this work, we present an initial framework for developing a DT, discuss and review the various challenges and considerations for DT deployment, and identify the opportunities that a DT can improve. The presented framework is intended to help developers formulate a strategy when approaching DT development for condition monitoring. In conclusion, a DT use case for a reactor coolant pump is presented to demonstrate the proposed framework.

advanced sensor instrumentation↗

Redesign of the Timeline Generator at Fermilab using a web-based Flutter application, GraphQL API and an IOC

The control system at Fermilab is undergoing an evolution with a shift towards web-based applications with connections to the EPICS infrastructure. The Timeline Generator (TLG) is an application that serves to coordinate events across the lab using different timing links. These links include the Tevatron clock (TCLK), a 10 MHz serial link with events encoded at 20Hz and Machine Data (MDAT), a communication link with states encoded at 720Hz. This paper covers the redesign of the major components of the TLG. This includes a web-based Flutter application for building timelines. A placement service is in use that has a GraphQL interface and uses a timeline input to compute a schedule of events and states. The Flutter application sends this computed schedule to the TLG IOC via a GraphQL interface to the Data Pool Manager (DPM). The TLG IOC runs on an Arria FPGA, the Accelerator Clock Generator (ACLK-GEN), which is responsible for writing the events and states on to the different timing links.

Carmichael, Linden [Fermilab]↗

CHARACTERIZING AND CONTROLLING RECOVERY AND RECRYSTALLIZATION IN NIOBIUM FOR IMPROVED SRF CAVITY PERFORMANCE

Crystal defects, such as dislocations and low-angle boundaries, provide sources of magnetic flux trapping in the Nb materials used for superconducting radio frequency (SRF) resonating cavities. Improving the performance of SRF cavities, as measured through the quality factor, requires reducing these defects. SRF cavity production involves deformation processing, such as rolling and forming, and strategic annealing heat treatments. The resulting microstructures can be recovered, recrystallized, or both. Because recovery leaves many defects that can trap flux, recrystallization should improve cavity performance. Thus, processing schedules that produce complete recrystallization without excessive grain growth need to be designed. Solutions to this problem require understanding physical metallurgy and differentiating between recovered and recrystallized regions of microstructure. Backscattered electron microscopy techniques are applied to this end. We demonstrate that the conditions required to produce fully recrystallized microstructures depend on Nb impurity content, suggesting that processing schedules may need to be adjusted by material heat or lot. We also demonstrate that processing can be used to control growth of recrystallized grains to maintain mechanical strength in fully recrystallized materials. Forming cavities from cold-rolled Nb sheet material may provide strategic new routes to obtain microstructures that improve SRF cavity performance.

Taleff, E. [The University of Texas at Austin]↗