Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “queuing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Queued Up: Characteristics of Power Plants Seeking Transmission Interconnection As of the End of 2022 [Slides]

Proposed large-scale electric generation and storage projects must apply for interconnection to the bulk power system via interconnection queues. While most projects that apply for interconnection are not subsequently built, data from these queues nonetheless provide a general indicator for mid-term trends in developer interest. Berkeley Lab compiled and analyzed data from all seven ISOs/RTOs in concert with 35 non-ISO utilities, representing an estimated 85% of all U.S. electricity load. We include all "active" projects in these generation interconnection queues through the end of 2022, as well as data on "operational" and "withdrawn" projects where those data are available. We find that the amount of new electric capacity in these queues is growing dramatically, with over 2,000 gigawatts (GW) of total generation and storage capacity now seeking connection to the grid (over 95% of which is for zero-carbon resources like solar, wind, and battery storage). Solar (947 GW) and battery storage (~680 GW) are – by far – the fastest growing resources in the queues; combined they accounted for over 80% of new capacity entering the queues in 2022. Substantial wind (300 GW) capacity is also seeking interconnection, 38% of which is for offshore projects (113 GW). In total, about 1,250 GW of zero-carbon generating capacity is currently seeking transmission access, as is 82 GW of natural gas capacity. Hybrids projects (co-locating multiple generation and/or storage types) comprise a large – and increasing – share of proposed projects, particularly in CAISO and the non-ISO West. 457 GW of solar hybrids (primarily solar+battery) and 24 GW of wind hybrids are currently active in the queues; over half of battery storage in the queues is paired with generation. However, much of this proposed capacity will be withdrawn from the queues and not built. Among a subset of queues for which data are available, only 21% of the projects (and 14% of capacity) seeking connection from 2000 to 2017 have been built as of the end of 2022. Additionally, interconnection wait times are on the rise: The typical duration from connection request to commercial operation increased from <2 years for projects built in 2000-2007 to nearly 4 years for those built in 2018-2022 (with a median of 5 years for projects built in 2022).

24 POWER TRANSMISSION AND DISTRIBUTION↗

Queued Up: 2024 Edition, Characteristics of Power Plants Seeking Transmission Interconnection As of the End of 2023 [Slides]

Electric transmission system operators (ISOs, RTOs, or utilities) require projects seeking to connect to the grid to undergo a series of impact studies before they can be built. This process establishes what new transmission equipment or upgrades may be needed before a project can connect to the system and assigns the costs of that equipment. The lists of projects in this process are known as “interconnection queues”. The amount of new electric capacity in these queues is growing dramatically, with nearly 2,600 gigawatts (GW) of total generation and storage capacity now seeking connection to the grid (over 95% of which is for zero-carbon resources like solar, wind, and battery storage). However, most projects that apply for interconnection are ultimately withdrawn, and those that are built are taking longer on average to complete the required studies and become operational. Data from these queues nonetheless provide a general indicator for mid-term trends in developer interest.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Queued Up: 2026 Edition, Characteristics of Power Plants Seeking Transmission Interconnection As of the End of 2025 [Slides]

Electric transmission system operators (ISOs, RTOs, or utilities) require proposed power plants seeking to connect to the transmission grid to undergo a series of impact studies before they can be built. This process establishes what new transmission equipment or upgrades may be needed before a project can connect to the system and assigns the costs of that equipment. The lists of projects in this process are known as “interconnection queues”. In collaboration with https://www.interconnection.fyi. Berkeley Lab compiled, aggregated, and cleaned interconnection queue data from >50 transmission grid operators (7 ISO/RTOs and 50 non-ISO balancing areas), which collectively represent ~98% of currently installed U.S. electric generating capacity.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Monitoring and queuing for sift

The SIFT instrumentation is called the "Window." This window was designed to collect internal data from SIFT while having minimal overhead. Window consists of Sender and Relay components. Sender is to be run on processors 0..5 and Relay will run on processor 6. Sender will gather values (currently 12) during the subframe allocated to a task and broadcast these values at the start of the next subframe. This timing was selected to guarantee Relay 3.2ms to collect and transmit the data.

Wilson, L.↗

A modeling framework for designing and evaluating curbside traffic management policies at Dallas-Fort Worth International Airport

Emerging mobility technologies are changing the transportation system landscape. This is especially evident at airports, such as the Dallas-Fort Worth International Airport (DFW). Without careful analysis, these changes could lead to inefficient and costly airport operations. This paper presents a modeling framework that integrates travel mode encoding, demand projection, and microsimulation to enable airports to develop, simulate, and evaluate curbside traffic managements policies and measure their impact. Here, the framework is utilized to analyze several traffic scenarios and policies for DFW: a baseline scenario which represents DFW traffic pattern as observed in 2018 and projected to 2045, a transit network company (TNC) electrification policy, a TNC queuing policy, a policy that increased transit ridership, a bus-only policy which considers the use of only buses inside DFW, an autonomous vehicle (AV) policy which investigates the impact of autonomous vehicle (AV) adoption on airport operations, and an example COVID-19 scenario which models the impact of the COVID19 pandemic. The simulations’ results demonstrate that: increasing the DFW transit ridership postpones the need for airport curbside expansion the most; encouraging shared-mobility with the bus-only policy produces the most savings in curbside congestion delays; automation and electrification for all passenger vehicle trips to/from DFW generates the most saving in fuel consumption and emissions; and uncontrolled AV adoption incurs the highest increase in fuel consumption, delay, and emissions and could require immediate airport capacity extension. Without policy intervention or investment in additional infrastructure capacity, these results predict the current operations would face significant congestion on high demand days starting as early as 2028. While derived in close partnership with DFW, the methodology presented here can be generalized to any airport.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Towards an Introspective Dynamic Model of Globally Distributed Computing Infrastructures

Large-scale scientific collaborations like ATLAS, Belle II, CMS, DUNE, and others involve hundreds of research institutes and thousands of researchers spread across the globe. These experiments generate petabytes of data, with volumes soon expected to reach exabytes. Consequently, there is a growing need for computation, including structured data processing from raw data to consumer-ready derived data, extensive Monte Carlo simulation campaigns, and a wide range of end-user analysis. To manage these computational and storage demands, centralized workflow and data management systems are implemented. However, decisions regarding data placement and payload allocation are often made disjointly and via heuristic means. A significant obstacle in adopting more effective heuristic or AI-driven solutions is the absence of a quick and reliable introspective dynamic model to evaluate and refine alternative approaches. In this study, we aim to develop such an interactive system using real-world data. By examining job execution records from the PanDA workflow management system, we have pinpointed key performance indicators such as queuing time, error rate, and the extent of remote data access. The dataset includes five months of activity. Additionally, we are creating a generative AI model to simulate time series of payloads, which incorporate visible features like category, event count, and submitting group, as well as hidden features like the total computational load—derived from existing PanDA records and computing site capabilities. These hidden features, which are not visible to job allocators, whether heuristic or AI-driven, influence factors such as queuing times and data movement.

kilic, Ozgur Ozan [Brookhaven National Laboratory ↗

The Empirical Effect of Fleet Optimization on Synchronization and Rebound Effects in Heat Pump Water Heaters

Demand response is a growing concept in light of the internet of things and an increasing need for grid flexibility. Water heaters are one of the preferred devices for providing demand response for grid services and peak management due to their capability to store energy. The efficient use of water heaters for demand response requires consideration of the associated load effects such as synchronization of device schedules and rebound effect. These effects present a significant challenge. Despite the importance of the mentioned effects for water heater queuing and scheduling, there has been no effort to quantify and empirically validate their impact. This study attempts to address this gap by offering two methods - Ward clustering and Euclidean K-means - to evaluate the extent of synchronization in a fleet of 42 water heaters in Atlanta, GA. Using the aforementioned methods on the measured data, we find evidence of convergence of water heater loads as a result of optimization compared to an idle period and analyzed their impact.

demand response↗

Adaptive job and resource management for the growing quantum cloud

As the popularity of quantum computing continues to grow, efficient quantum machine access over the cloud is critical to both academic and industry researchers across the globe. And as cloud quantum computing demands increase exponentially, the analysis of resource consumption and execution characteristics are key to efficient management of jobs and resources at both the vendor-end as well as the client-end. While the analysis and optimization of job / resource consumption and management are popular in the classical HPC domain, it is severely lacking for more nascent technology like quantum computing.This paper proposes optimized adaptive job scheduling to the quantum cloud taking note of primary characteristics such as queuing times and fidelity trends across machines, as well as other characteristics such as quality of service guarantees and machine calibration constraints. Key components of the proposal include a) a prediction model which predicts fidelity trends across machine based on compiled circuit features such as circuit depth and different forms of errors, as well as b) queuing time prediction for each machine based on execution time estimations. Altogether, this proposal is evaluated on simulated IBM machines across a diverse set of quantum applications and system loading scenarios, and is able to reduce wait times by over 3x and improve fidelity by over 40% on specific usecases, when compared to traditional job schedulers.

97 MATHEMATICS AND COMPUTING↗

In-Situ Calibrated Digital Process Twin Models for Resource Efficient Manufacturing

The chief objective of manufacturing process improvement efforts is to significantly minimize process resources such as time, cost, waste, and consumed energy while improving product quality and process productivity. This paper presents a novel physics-informed optimization approach based on artificial intelligence (AI) to generate digital process twins (DPTs). The utility of the DPT approach is demonstrated in the case of finish machining of aerospace components made from gamma titanium aluminide alloy (γ-TiAl). This particular component has been plagued with persistent quality defects, including surface and sub-surface cracks, which adversely affect resource efficiency. Previous process improvement efforts have been restricted to anecdotal post-mortem investigation and empirical modeling, which fail to address the fundamental issue of how and when cracks occur during cutting. In this work, the integration of in-situ process characterization with modular physics-based models is presented, and machine learning algorithms are used to create a DPT capable of reducing environmental and energy impacts while significantly increasing yield and profitability. Based on the preliminary results presented here, we report an improvement in the overall embodied energy efficiency of over 84%, 93% in process queuing time, 2% in scrap cost, and 93% in queuing cost has been realized for γ-TiAl machining using our novel approach.

42 ENGINEERING↗

Cost performance satellite design using queueing theory

A modified Poisson arrival, infinite server queuing model is used to determine the effects of limiting the number of broadcast channels (C) of a direct broadcast satellite used for public service purposes (remote health care, education, etc.). The model is based on the reproductive property of the Poisson distribution. A difference equation has been developed to describe the change in the Poisson parameter. When all initially delayed arrivals reenter the system a (C plus 1) order polynomial must be solved to determine the effective value of the Poisson parameter. When less than 100% of the arrivals reenter the system the effective value must be determined by solving a transcendental equation. The model was used to determine the minimum number of channels required for a disaster warning satellite without degradation in performance. Results predicted by the queuing model were compared with the results of digital simulation.

Hein, G. F.↗

Modeling and measurement of fault-tolerant multiprocessors

The workload effects on computer performance are addressed first for a highly reliable unibus multiprocessor used in real-time control. As an approach to studing these effects, a modified Stochastic Petri Net (SPN) is used to describe the synchronous operation of the multiprocessor system. From this model the vital components affecting performance can be determined. However, because of the complexity in solving the modified SPN, a simpler model, i.e., a closed priority queuing network, is constructed that represents the same critical aspects. The use of this model for a specific application requires the partitioning of the workload into job classes. It is shown that the steady state solution of the queuing model directly produces useful results. The use of this model in evaluating an existing system, the Fault Tolerant Multiprocessor (FTMP) at the NASA AIRLAB, is outlined with some experimental results. Also addressed is the technique of measuring fault latency, an important microscopic system parameter. Most related works have assumed no or a negligible fault latency and then performed approximate analyses. To eliminate this deficiency, a new methodology for indirectly measuring fault latency is presented.

Shin, K. G.↗

Performance analysis of fault-tolerant systems in parallel execution of conversations

The execution overhead inherent in the conversation scheme, which is a scheme for realizing fault-tolerant cooperating processes free of the domino effect, is analyzed. Multiprocessor/multicomputer systems capable of parallel execution of conversation components are considered and a queuing network model of such systems is adopted. Based on the queuing model, various performance indicators, including system throughput, average number of processors idling inside a conversation due to the synchronization required, and average time spent in the conversation, have been evaluated numerically for several application environments. The numeric results are discussed and several essential performance characteristics of the conversation scheme are derived. For example, when the number of participant processes is not large, say less than six, the system performance is highly affected by the synchronization required on the processes in a conversation, and not so much by the probability of acceptance-test failure.

Kim, K. H.↗

Average waiting time in FDDI networks with local priorities

A method is introduced to compute the average queuing delay experienced by different priority group messages in an FDDI node. It is assumed that no FDDI MAC layer priorities are used. Instead, a priority structure is introduced to the messages at a higher protocol layer (e.g. network layer) locally. Such a method was planned to be used in Space Station Freedom FDDI network. Conservation of the average waiting time is used as the key concept in computing average queuing delays. It is shown that local priority assignments are feasable specially when the traffic distribution is asymmetric in the FDDI network.

Gercek, Gokhan↗

Job Scheduling Under the Portable Batch System

The typical batch queuing system schedules jobs for execution by a set of queue controls. The controls determine from which queues jobs may be selected. Within the queue, jobs are ordered first-in, first-run. This limits the set of scheduling policies available to a site. The Portable Batch System removes this limitation by providing an external scheduling module. This separate program has full knowledge of the available queued jobs, running jobs, and system resource usage. Sites are able to implement any policy expressible in one of several procedural language. Policies may range from "bet fit" to "fair share" to purely political. Scheduling decisions can be made over the full set of jobs regardless of queue or order. The scheduling policy can be changed to fit a wide variety of computing environments and scheduling goals. This is demonstrated by the use of PBS on an IBM SP-2 system at NASA Ames.

Henderson, Robert L.↗

CCSDS Advanced Orbiting Systems Virtual Channel Access Service for QoS MACHETE Model

To support various communications requirements imposed by different missions, interplanetary communication protocols need to be designed, validated, and evaluated carefully. Multimission Advanced Communications Hybrid Environment for Test and Evaluation (MACHETE), described in "Simulator of Space Communication Networks" (NPO-41373), NASA Tech Briefs, Vol. 29, No. 8 (August 2005), p. 44, combines various tools for simulation and performance analysis of space networks. The MACHETE environment supports orbital analysis, link budget analysis, communications network simulations, and hardware-in-the-loop testing. By building abstract behavioral models of network protocols, one can validate performance after identifying the appropriate metrics of interest. The innovators have extended the MACHETE model library to include a generic link-layer Virtual Channel (VC) model supporting quality-of-service (QoS) controls based on IP streams. The main purpose of this generic Virtual Channel model addition was to interface fine-grain flow-based QoS (quality of service) between the network and MAC layers of the QualNet simulator, a commercial component of MACHETE. This software model adds the capability of mapping IP streams, based on header fields, to virtual channel numbers, allowing extended QoS handling at link layer. This feature further refines the QoS v existing at the network layer. QoS at the network layer (e.g. diffserv) supports few QoS classes, so data from one class will be aggregated together; differentiating between flows internal to a class/priority is not supported. By adding QoS classification capability between network and MAC layers through VC, one maps multiple VCs onto the same physical link. Users then specify different VC weights, and different queuing and scheduling policies at the link layer. This VC model supports system performance analysis of various virtual channel link-layer QoS queuing schemes independent of the network-layer QoS systems.

Jennings, Esther H.↗

Space Link Extension Protocol Emulation for High-Throughput, High-Latency Network Connections

New space missions require higher data rates and new protocols to meet these requirements. These high data rate space communication links push the limitations of not only the space communication links, but of the ground communication networks and protocols which forward user data to remote ground stations (GS) for transmission. The Consultative Committee for Space Data Systems, (CCSDS) Space Link Extension (SLE) standard protocol is one protocol that has been proposed for use by the NASA Space Network (SN) Ground Segment Sustainment (SGSS) program. New protocol implementations must be carefully tested to ensure that they provide the required functionality, especially because of the remote nature of spacecraft. The SLE protocol standard has been tested in the NASA Glenn Research Center's SCENIC Emulation Lab in order to observe its operation under realistic network delay conditions. More specifically, the delay between then NASA Integrated Services Network (NISN) and spacecraft has been emulated. The round trip time (RTT) delay for the continental NISN network has been shown to be up to 120ms; as such the SLE protocol was tested with network delays ranging from 0ms to 200ms. Both a base network condition and an SLE connection were tested with these RTT delays, and the reaction of both network tests to the delay conditions were recorded. Throughput for both of these links was set at 1.2Gbps. The results will show that, in the presence of realistic network delay, the SLE link throughput is significantly reduced while the base network throughput however remained at the 1.2Gbps specification. The decrease in SLE throughput has been attributed to the implementation's use of blocking calls. The decrease in throughput is not acceptable for high data rate links, as the link requires constant data a flow in order for spacecraft and ground radios to stay synchronized, unless significant data is queued a the ground station. In cases where queuing the data is not an option, such as during real time transmissions, the SLE implementation cannot support high data rate communication.

Computer Networking↗

Queue wait time prediction in high performance computing (HPC) systems

High Performance Computing (HPC) systems are critical enablers for groundbreaking scientific research across various domains. Efficient resource allocation, facilitated by job scheduling, is paramount for maximizing the utilization of HPC systems. However, the variability in wait times for queued jobs poses challenges for users, necessitating accurate job wait time estimation. This paper explores the influence of job characteristics, including job size (the number of nodes requested and walltime), the queue to which the job is submitted and other resource requirements, on job wait times in leadership-class HPC systems. Focusing on the Theta Cray XC40 and Polaris machines at Argonne National Laboratory, the study evaluates the performance of different supervised learning algorithms in predicting job wait times. It also evaluates the impact of data preprocessing, including outlier detection, Principal Component Analysis (PCA), and feature selection, on the performance of wait time prediction models. The findings reveal insights into the relationship between job characteristics and wait times, offering a foundation for optimizing resource allocation and enhancing user experience. The methodologies and tools developed in this study are adaptable to other leadership-class HPC systems, providing a valuable contribution to the broader HPC community aiming to improve job scheduling efficiency and user satisfaction.

Okafor, Nwamaka↗

Enabling machine learning-ready HPC ensembles with Merlin

With the growing complexity of computational and experimental facilities, many scientific researchers are turning to machine learning (ML) techniques to analyze large scale ensemble data. With complexities such as multi-component workflows, heterogeneous machine architectures, parallel file systems, and batch scheduling, care must be taken to facilitate this analysis in a high performance computing (HPC) environment. Here, we present Merlin, a workflow framework to enable large ML-friendly ensembles of scientific HPC simulations. By augmenting traditional HPC with distributed compute technologies, Merlin aims to lower the barrier for scientific subject matter experts to incorporate ML into their analysis. As a producer–consumer workflow model, Merlin enables multi-machine, cross-batch job, dynamically allocated yet persistent workflows capable of utilizing surge-compute resources. Key features of Merlin are a flexible HPC-centric interface, low per-task overhead, multi-tiered fault recovery, and a hierarchical sampling algorithm that allows for $\mathscr{O}$(N) task execution and $\mathscr{O}$(N ln N) task queuing to ensembles of millions of tasks. In addition to Merlin’s design, we test the algorithm’s performance in an HPC center and demonstrate the ability to enqueue 40 million simulations in 100 s, with a 30 millisecond per-task overhead that is independent of ensemble size. Finally, we describe some example applications that Merlin has enabled on leadership-class HPC resources, such as the ML-augmented optimization of nuclear fusion experiments and the calibration of infectious disease models to study the progression of and possible mitigation strategies for COVID-19.

97 MATHEMATICS AND COMPUTING↗