Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “cluster scheduling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Detailed Skylab ECS consumables analysis for the interim revision flight plan (November, 1972, SL-1 launch)

The consumables analysis was performed for the Skylab 2, 3, and 4 Preliminary Reference Interim Revision Flight Plan. The analysis and the results are based on the mission requirements as specified in the flight plan and on other available data. The results indicate that the consumables requirements for the Skylab missions allow for remaining margins (percent) of oxygen, nitrogen, and water nominal as follows: 83.5, 90.8, and 88.7 for mission SL-2; 57.1, 64.1, and 67.3 for SL-3; and 30.8, 44.3, and 46.5 for SL-4. Performance of experiment M509 as scheduled in the flight plan results in venting overboard the cluster atmosphere. This is due to the addition of nitrogen for propulsion and to the additional oxygen introduced into the cabin when the experiment is performed with the crewman suited.

Wells, C.↗

Space shuttle with common fuel tank for liquid rocket booster and main engines (supertanker space shuttle)

An operation and schedule enhancement is shown that replaces the four-body cluster (Space Shuttle Orbiter (SSO), external tank, and two solid rocket boosters) with a simpler two-body cluster (SSO and liquid rocket booster/external tank). At staging velocity, the booster unit (liquid-fueled booster engines and vehicle support structure) is jettisoned while the remaining SSO and supertank continues on to orbit. The simpler two-bodied cluster reduces the processing and stack time until SSO mate from 57 days (for the solid rocket booster) to 20 days (for the liquid rocket booster). The areas in which liquid booster systems are superior to solid rocket boosters are discussed. Alternative and future generation vehicles are reviewed to reveal greater performance and operations enhancements with more modifications to the current methods of propulsion design philosophy, e.g., combined cycle engines, and concentric propellant tanks.

Thorpe, Douglas G.↗

Enhancing Active Distribution Systems Resilience by Fully Distributed Self-Healing Strategy

Distributed restoration can exploit smart grid technologies to enhance the resilience of active distribution networks toward a self-healing smart grid. However, the large number of decision variables, especially the binary ones for reconfiguration, bring challenges to developing scalable distributed distribution service restoration (DDSR) strategies. This paper proposes a fully distributed solution procedure based on the alternating direction method of multipliers (ADMM) for mixed-integer programming problems and applies to develop the DDSR framework. The method consists of relax-drive-polish phases, 1) relaxing binary variables, and applying the convex ADMM as a warm start; 2) driving the solutions toward Boolean values through a proximal operator; 3) fixing the obtained binding binary variables and solving the rest of the problem to polish results and achieve a high-quality suboptimal solution. Then, an autonomous clustering strategy and consensus ADMM are integrated with the proposed method to realize the fully distributed cluster-based framework of DDSR. This framework can first determine DER scheduling and switch status for reconfiguration to energize the out-of-service areas from local faults, and then provide the load restoration solution in a distributed manner for total blackouts in large-scale distribution networks. Furthermore, the effectiveness and scalability of the proposed DDSR framework are demonstrated through testing on the IEEE 123-node, IEEE 8500-node, and synthetic 100k-node test feeders.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Informed Feature Selection for Data Clustering of CSP Plant Production

To make concentrating solar power (CSP) more cost competitive, rigourous optimizations must be run to improve plant design and operations. However, these optimizaitons rely on time consuming annual simulations that solve an electricity dispatch scheduling problem to maximize plant revenue. To reduce the runtime of annual dispatch simulations of CSP plants, a data clustering approach is utilized. This approach assumes that like days of revenue and electricity generation can be identified using weather and price data. Although weather and price are important factors for electricity production, this work investigates how thermal energy storage (TES) inventory at the beginning of a day, denoted as Si, can be used as a supplemental feature to group like days. A framework for creating and training a deep neural network to predict Si is proposed. This model is validated and assessed using eleven sets of testing data that were not used during training. Then, the data clustering approach is performed three seperate times with features of weather and price along with either Si from the neural network, Si from the full annual simulation, or no Si. Ultimately, the results suggest that using Si as an additional clustering feature improves the data clustering simulation accuracy by 1.4%.

Tuman, Matthew J. (ORCID:000900038772051X)↗

SchedInspector: A Batch Job Scheduling Inspector Using Reinforcement Learning

Improving the performance of job executions is an important goal of HPC batch job schedulers, such as minimizing job waiting time, slowdown, or completion time. Such a goal is often accomplished using carefully designed heuristics based on job features, such as job size and job duration. However, these heuristics overlook important runtime factors (e.g., cluster availability and waiting job patterns), which may vary across time and make a previously sound scheduling decision not hold any longer. In this study, we propose a new approach to incorporate runtime factors into batch job scheduling for better job execution performance. The key idea is to add a scheduling inspector on top of the base job scheduler to scrutinize its scheduling decisions. The inspector will take the runtime factors into consideration and accordingly determine the fitness of the scheduled job. It then either accepts the scheduled job or rejects it and asks the base schedulers to try again later. We realize such an inspector, namely SchedInspector, by leveraging the intelligence of reinforcement learning. Through extensive experiments, we show SchedInspector can intelligently integrate the runtime factors into various batch job scheduling policies, including the state-of-the-art one, to gain better job execution performance, such as smaller average bounded job slowdown (up to 69% better) or average job waiting time (up to 52% better), across various real-world workloads. We also show that although rejecting scheduling decisions may leave the resources idle hence affect the system utilization, SchedInspector is able to achieve the job execution performance improvement with marginal impact on the system utilization (typically less than 1%). We consider one key advantage of SchedInspector is it automatically learns to work with and improve existing job scheduling policies without changing them, which makes it promising to serve as a generic enhancer for various batch job scheduling policies.

Zhang, Di↗

Job Management Requirements for NAS Parallel Systems and Clusters

A job management system is a critical component of a production supercomputing environment, permitting oversubscribed resources to be shared fairly and efficiently. Job management systems that were originally designed for traditional vector supercomputers are not appropriate for the distributed-memory parallel supercomputers that are becoming increasingly important in the high performance computing industry. Newer job management systems offer new functionality but do not solve fundamental problems. We address some of the main issues in resource allocation and job scheduling we have encountered on two parallel computers - a 160-node IBM SP2 and a cluster of 20 high performance workstations located at the Numerical Aerodynamic Simulation facility. We describe the requirements for resource allocation and job management that are necessary to provide a production supercomputing environment on these machines, prioritizing according to difficulty and importance, and advocating a return to fundamental issues.

Saphir, William↗

Extension of constrained incremental Newton-Raphson scheme to generalized loading fields

This paper develops numerical strategies which enable the constrained incremental Newton-Raphson scheme to handle the static response of structure to loading fields with completely generalized histories. This is made possible through the use of specially warped hyperelliptic constraint surfaces which control successive or clustered load steps in the vicinity of loading events with specific timing schedules. Such an approach enables improved convergence and stability characteristics. Due to the generality of the methodology, pre- and postbuckling behavior caused by both kinematic and material nonlinearity can be handled. To demonstrate the scheme, the results of several bench-mark problems are also presented. These include situations involving nonlinear kinematics as well as highly history-dependent elastic-plastic and thermoelastic-plastic material behavior.

Padovan, J.↗

Ly alpha and IR galaxy companions of high redshift damped Ly alpha QSO absorbers

We have used a Near-Infrared Camera and Multi-Object Spectrometer (NICMOS3) HgCdTe 256x256 array detector with the Infrared (IR) camera on the 2.3m telescope at Steward Observatory to image several Quasi-Stellar Object (QSO) fields. The limiting magnitude is K'(2.1 microns) = 21.0 - 21.5 mag per square arcsec for a 3 sigma detection in 3 hours of in-field chopping observations. Each QSO line-of-sight samples several known absorbers with Mg2(lambda)2796-2803 A and/or C4(lambda)1548-1551 A absorption doublets. The equivalent width distributions of the low and high ionization absorption lines of the absorber sample are identical to those of the parent population of all absorbers. This selection process, used already for a spectroscopic survey of Mg2 absorption lines in C4-selected absorption systems at high z, gives a methodical approach to observing, reduces the observer biases, and makes a more efficient use of telescope time. This selection guarantees that imaging of the sample of QSO fields will provide complete sampling of the whole population of high z QSO absorbers. Follow-up optical and IR spectroscopy of these objects is scheduled for redshift measurement and confirmation of the absorbing galaxies and the cluster members.

Caulet, Adeline↗

Preventive Power Outage Estimation Based on a Novel Scenario Clustering Strategy

The increasing occurrence of extreme weather events is challenging power grid operation. For extreme weather events, the system operator is responsible for estimating the power outages and scheduling the restoration resources. This paper proposes an outage evaluation framework to identify the possible unserved load profiles, vulnerable areas, and mobile energy adequacy. The outputs of an outage prediction model tool are used to generate numerous faulted line scenarios. Next, each scenario's nodal unserved load profile is obtained by solving a three-phase restoration model that considers repair crews and mobile energy resources (MERs). Then, a novel scenario clustering strategy is developed to cluster the unserved load profiles into multiple representative profiles which the system operator can focus on. Finally, case studies on a distribution system evaluate the damage caused by an extreme weather event and verify the effectiveness of the proposed scenario clustering strategy.

MATHEMATICS AND COMPUTING,POWER TRANSMISSION AND D↗

A Low-Cost Clustered Archive Approach for Storing Remote Sensing Data

As part of NASA's Earth Observing System (EOS) Data and Information System (EOSDIS), the Moderate Resolution Imaging Spectrometer (MODIS) Data Processing System (MODAPS) is now processing data from two instruments on the EOS flag ship spacecraft Terra and Aqua. Between the two, MODAPS is generating over 1 Terrabyte of data per day and has surpassed 3 Petabytes of total data. The bulk of the data is stored near-line in StorageTek Powerhorn tape jukeboxes. Accessing data that has been moved to tape involves submitting an order, scheduling the tape, and waiting for the data to become available. I am developing a low cost clustered archive that could enable storing a very large amount of data such as the MODIS data described above in an organized fashion on a cluster of commodity hardware using low cost SATA hard drives such that the files are directly available online. The system takes full advantage of Open Source software, using the GNULinux operating system, PostgreSQL relational database and the Apache HTTP Server. This poster session will depict my approach, the interface to the archive, a brief discussion of the internals of the system and some performance numbers from my prototyping. I will also describe various costs and benefits of this approach versus the traditional large tape jukebox approach currently in use in the EOSDIS.

Tilmes, Curt↗

Maskman

SAND2025-04369O Maskman is a user-friendly tool designed to create hex masks, which are essential for optimizing application performance in high-performance computing environments. By converting a list of integers into binary and then hex masks, Maskman simplifies the process of setting application affinity. This ensures that software runs efficiently on specific nodes within a computing cluster. Ideal for researchers and developers, Maskman streamlines the preparation of inputs for HPC schedulers, enhancing resource management and improving overall system performance. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Pase, Douglas [Sandia National Lab. (SNL-CA), Live↗

Resource Selection Using Execution and Queue Wait Time Predictions

We developed techniques to predict application execution times for instance-based learning with an average error of 33% of average run time. We developed techniques to predict queue wait times that included a simulation of scheduling algorithms and execution time predictions. We implemented these techniques for the NAS Origin cluster.

Smith, Warren↗

Advanced rocket propulsion

Existing NASA research contracts are supporting development of advanced reinforced polymer and metal matrix composites for use in liquid rocket engines of the future. Advanced rocket propulsion concepts, such as modular platelet engines, dual-fuel dual-expander engines, and variable mixture ratio engines, require advanced materials and structures to reduce overall vehicle weight as well as address specific propulsion system problems related to elevated operating temperatures, new engine components, and unique operating processes. High performance propulsion systems with improved manufacturability and maintainability are needed for single stage to orbit vehicles and other high performance mission applications. One way to satisfy these needs is to develop a small engine which can be clustered in modules to provide required levels of total thrust. This approach should reduce development schedule and cost requirements by lowering hardware lead times and permitting the use of existing test facilities. Modular engines should also reduce operational costs associated with maintenance and parts inventories.

Obrien, Charles J.↗

A Study of Parallel Scalability and Dynamic Workload Balancing in GlennICE

The Glenn Icing Computational Environment (GlennICE) is a computational tool designed to calculate ice growth on complex three-dimensional geometries. It utilizes user-supplied computational fluid dynamics solutions for the geometry of interest. Key developments include advancements in convergence of collection efficiency, trajectory optimization, and refinement methodology. These improvements have significantly enhanced GlennICE’s efficiency for practical engineering applications. A recent study focused on benchmarking GlennICE’s scalability in a parallel environment using static scheduling. Findings indicated a potential twofold increase in efficiency through workload balance enhancements. This paper presents an analysis of the solver’s new workload balancing improvements, incorporating shared memory and dynamic scheduling routines. Results demonstrate a highly efficient and consistent algorithm across high-performance computing clusters.

Computational Icing↗

A Study of Parallel Scalability and Dynamic Workload Balancing in GlennICE

The Glenn Icing Computational Environment (GlennICE) is a computational tool designed to calculate ice growth on complex three-dimensional geometries. It utilizes user-supplied computational fluid dynamics solutions for the geometry of interest. Key developments include advancements in convergence of collection efficiency, trajectory optimization, and refinement methodology. These improvements have significantly enhanced GlennICE’s efficiency for practical engineering applications. A recent study focused on benchmarking GlennICE’s scalability in a parallel environment using static scheduling. Findings indicated a potential twofold increase in efficiency through workload balance enhancements. This paper presents an analysis of the solver’s new workload balancing improvements, incorporating shared memory and dynamic scheduling routines. Results demonstrate a highly efficient and consistent algorithm across high-performance computing clusters.

Computational Icing↗

VC3: Virtual Clusters for Community Computation

A traditional HPC computing facility provides a large amount of computing power but has a fixed environment designed to satisfy local needs. This makes it very challenging for users to deploy complex applications that span multiple sites and require specific application software, scheduling middleware, or sharing policies. This project addressed many of these challenges by making it possible for researchers to easily aggregate and share resources, install custom software environments, and deploy clustering frameworks across multiple HPC facilities through the concept of “virtual clusters”. We designed and implemented a prototype virtual cluster facility that enabled unprivileged users to create dynamic aggregations of computing power across multiple sites, deployed with custom middleware and complex software dependencies. This service is hosted at the University of Chicago and available through the site virtualclusters.org.

97 MATHEMATICS AND COMPUTING↗

OctoFAS: A Two-Level Fair Scheduler That Increases Fairness in Network-Based Key-Value Storage

We identified a fairness problem in a network-based key-value storage system using Intel Storage Performance Development Kit (SPDK) in a multitenant environment. In such an environment, each tenant’s I/O service rate is not fairly guaranteed compared to that of other tenants. To address the fairness problem, we propose OctoFAS, a two-level fair scheduler designed to improve overall throughput and fairness among tenants. The two-level scheduler of OctoFAS consists of (i) inter-core scheduling and (ii) intra-core scheduling. Through inter-core scheduling, OctoFAS addresses the load imbalance problem that is inherent in SPDK on the storage server by dynamically migrating I/O requests from overloaded cores to underloaded cores, thereby increasing overall throughput. Intra-core scheduling prioritizes handling requests from starving tenants over well-fed tenants within core-specific event queues to ensure fair I/O services among multiple tenants. OctoFAS is deployed on a Linux cluster with SPDK. Through extensive evaluations, we found that OctoFAS ensures that the total system throughput remains high and balanced, while enhancing fairness by approximately 10% compared to the baseline, when both scheduling levels operate in a hybrid fashion.

97 MATHEMATICS AND COMPUTING↗