Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “cluster scheduling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

68 records · Page 4

Decentralized Failure-Tolerant Optimization of Electric Vehicle Charging

We present a decentralized failure-tolerant algorithm for optimizing electric vehicle (EV) charging, using charging stations as computing agents. The algorithm is based on the alternating direction method of multipliers (ADMM) and it has the following features: (i) It handles capacity, peak demand, and ancillary services coupling constraints. (ii) It does not require a central agent collecting information and performing coordination (e.g. an aggregator), instead all agents exchange information and computations are carried out in a fully decentralized fashion. (iii) It can withstand the failure of any number of computing agents, as long as the remaining computing agents are in a connected communications network. We construct this algorithm by reformulating the optimal EV charging problem in a decomposable form, amenable to ADMM, and then developing efficient decentralized solution methods for the subproblems dealing with coupling constraints. We conduct numerical experiments on industry-scale synthetic EV charging datasets, with up to 1,152 charging stations, using a high performance computing cluster. The experiments demonstrate that the proposed algorithm can solve the optimal EV charging problem fast enough to permit the integration of EV charging with real-time electricity markets, even in the presence of failures.

42 ENGINEERING↗

Plentiful electricity turns wholesale prices negative

In 2020, average wholesale electricity prices in the United States fell to $21/MWh, their lowest level since the beginning of the 21st century. Low natural gas prices and the proliferation of low marginal cost resources like wind and solar had already established a trend toward lower wholesale prices, and this trend was exacerbated by declining electricity demand due to the Covid-19 pandemic in 2020. Negative real-time hourly wholesale prices occurred in about 4% of all hours and wholesale market nodes across the United States, but these were not distributed evenly. Regional clusters emerged, for example, in the Permian Basin in western Texas, and in Kansas and western Oklahoma in the Southwest Power Pool (SPP), negative prices accounted for more than 25% of all hours. Negative electricity prices result either from local congestion of the transmission system leading supply to exceed demand locally or due to system-wide oversupply. Looking at the latter condition in SPP, we find that all major generator types contribute to this excess supply, because of limited ramping flexibility or self-scheduled out-of-market unit commitments. Additional monetary production incentives such as renewable energy credits or tax credits also enable negative bids; indeed, negative prices predominantly occur when demand levels are low and wind production levels are high. Frequent negative prices can inform the value of additional renewable energy investments at specific locations, the need for transmission and storage development, and opportunities load growth or adaptation.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Scalable, In-situ Data Clustering Data Analysis for Extreme Scale Scientific Computing (Final Report)

The objective of this project is to address challenges in the design and development of scalable in-situ data clustering and analytics algorithms and software. Our goal is to develop parallel software consisting of a set of spatio-temporal data clustering and anomaly detection functions, both of which are very important for large-scale analysis and have wide applicability for in-situ runs as well as post-processing analysis. Our design principles for in-situ analysis consider the following: (1) identify parts of the computation can be done close to the data within the nodes, while it is still in memory; (2) extract analysis components can (and should) be performed in remote staging and analysis nodes; (3) develop error-bound approximation methods for applications tolerable for small errors; (4) identify the type of derived distributions and statistics, for spatio-temporal data, that can be kept locally in order to both accelerate computations and meet energy constraints in subsequent iterations and phases; (5) use a self-describing data format so that data can be consistent and understood among local storage (memory and SSDs) and at staging and analysis nodes, thereby providing portability and flexibility; (6) develop service-oriented functions that can schedule in-situ and post-hoc analysis tasks based on the dynamic requirements of applications. Our development focus is to produce the parallel data analysis software/library that will be scalable, reusable, extensible, and generic for applications in different disciplines. The software will be able to run in-situ with the simulations as well as post-hoc analysis. This approach will satisfy many synergistic requirements for data intensive applications executed on data coming from instruments and experiments. In particular, the proposed multilevel approach is directly applicable to perform design tradeoffs for running part of the algorithms near the instruments and the rest on remote (analysis) systems.

97 MATHEMATICS AND COMPUTING↗

From Cell to System: Accelerated hpc Simulations of BESS Aging under Frequency Regulation and Arbitrage use cases

Lithium-ion battery energy storage systems (BESS) packs have emerged as a leading solution for grid-scale energy storage, enhancing resiliency and balancing load fluctuations. Yet, experimental characterization of large-format LIB packs-particularly to assess performance and degradation over hundreds of cycles - demands substantial hardware investment and multi-year testing campaigns. In this work, we couple a hierarchical, physics-based modeling framework agnostic to electrode chemistries with high-performance computing to accelerate systems level evaluation by upto two orders of magnitude. Building on the open-source liionpack platform, we implement cell, module, and pack-scale electrochemical models enriched with mechanistic aging mechanisms and deploy them on an HPC cluster to simulate 150−200kWh systems over 500 - 1,000 cycles with in days. We subject these virtual B ESS to both constant-current cycling and realistic grid service profiles spanning frequency regulation, ramp-rate support, and energy arbitrage-and quantify the resulting degradation patterns. Our results reveal that localized cell aging can induce substantial nonuniformity at module and pack levels, with service-specific cycling protocols driving distinct aging modes. This rapid, multiscale modeling approach provides a powerful design-space exploration tool for optimizing electrical architecture, control strategies, and operational schedules to prolong pack lifetime and lower total cost of ownership.

Ayalasomayajula, Surya [ORNL] (ORCID:0009000860788↗

HPC ODA Commons [SWR-26-003]

HPC ODA Commons is a community-driven platform for standardizing HPC operational data analytics. HPC sites generate enormous volumes of operational data - scheduler logs, accounting records, monitoring streams - but turning that data into actionable insight is needlessly hard. Each site builds bespoke parsers, schemas, and evaluation pipelines. Results can't be compared across institutions. Promising analytics ideas stay siloed because there's no shared language for describing the data, the experiments, or the outcomes. HPC ODA Commons fixes this by establishing community-governed contracts - versioned schemas, canonical artifacts, and benchmark recipes - that make ODA workflows discoverable, reproducible, and comparable. It pairs these standards with a practical, CLI-first toolkit that lets operators and researchers go from raw logs to standardized results without sending data off-cluster.

Menear, Kevin [National Laboratory of the Rockies ↗

Unsupervised Detection of SOC Spoofing in OCPP 2.0.1 EV Charging Communication Protocol Using One-Class SVM

The electric vehicles (EVs) market keeps growing globally; thus, it is critical to secure the EV charging communication protocols in order to guarantee reliable and fair charging operations among the customers. The Open Charge Point Protocol (OCPP) 2.0.1 supports the communication between the Electric Vehicle Supply Equipment (EVSE) and Charging Station Management Systems (CSMSs); therefore, it becomes vulnerable to several types of attacks, which aim to jeopardize smart charging, billing, and energy management. Specifically, OCPP 2.0.1 allows the self-reporting of the State of Charge (SOC) values, which makes it vulnerable to spoofing-based cyberattacks, which target manipulating the scheduling priorities, distorting the load forecasts, and extending the charging sessions in an unfair manner. In this paper, we try to address this type of attack by providing a comprehensive analysis of the SOC spoofing attacks and introducing a novel unsupervised detection framework based on the One-Class Support Vector Machine (OCSVM) algorithm. Specifically, two types of attack scenarios are analyzed (i.e., priority manipulation and session extension) by deriving engineered features that capture the nonlinear relationships under normal charging behavior. Detailed simulation-based results are derived by utilizing the DESL-EPFL Level 3 EV charging dataset. Our results demonstrate high F1-score and recall in identifying spoofed SOC values and that the proposed OCSVM model demonstrates superior performance compared to alternative clustering and deep-learning based detectors.

EV charging↗

Reaching new peaks for the future of the CMS HTCondor Global Pool

The CMS experiment at CERN employs a distributed computing infrastructure to satisfy its data processing and simulation needs. The CMS Submission Infrastructure team manages a dynamic HTCondor pool, aggregating mainly Grid clusters worldwide, but also HPC, Cloud and opportunistic resources. This CMS Global Pool, which currently involves over 70 computing sites worldwide and peaks at 350k CPU cores, is employed to successfully manage the simultaneous execution of up to 150k tasks. While the present infrastructure is sufficient to harness the current computing power scales, CMS latest estimates predict a noticeable expansion in the amount of CPU that will be required in order to cope with the massive data increase of the High-Luminosity LHC (HL-LHC) era, planned to start in 2027. This contribution presents the latest results of the CMS Submission Infrastructure team in exploring and expanding the scalability reach of our Global Pool, in order to preventively detect and overcome any barriers in relation to the HL-LHC goals, while maintaining high effciency in our workload scheduling and resource utilization.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Exploratory analysis and performance prediction of big data transfer in High-performance Networks

Big data transfer in large-scale scientific and business applications is increasingly carried out over connections with guaranteed bandwidth provisioned in High-performance Networks (HPNs) via advance bandwidth reservation. Provisioning agents need to carefully schedule data transfer requests, compute network paths, and allocate appropriate bandwidths. Such reserved bandwidths, if not fully utilized, could be simply wasted due to the exclusive access during the approved time window, and cause extra overhead and complexity for resource management. This calls for accurate performance prediction to reserve bandwidths that match actual needs and avoid over-provisioning. We employ machine learning algorithms to predict big data transfer performance based on extensive performance measurements collected in the past several years from data transfer tests using different protocols and toolkits between various end sites on several real-life physical or emulated testbeds. We first analyze the performance patterns in response to a comprehensive list of parameters in end-host systems, network connections, and data transfer applications, which motivate the use of machine learning and also help us identify the effects of latent factors. We then propose threshold- and clustering-based methods to eliminate negative effects of latent factors in data preprocessing and build a robust performance predictor based on customized domain-oriented loss functions. The performance of the proposed methods is verified by extensive experiments using SVR and RFR as well as theoretical analysis of the general performance bound.

97 MATHEMATICS AND COMPUTING↗

Conquering Data Chaos: Research Data Management with Kubernetes

Managing massive volumes of data and effectively making it accessible to researchers poses significant challenges and is a barrier to scientific discovery. In many cases, critical data is locked up in unwieldy file formats or one-off databases and is too large to effectively process on a single machine. This talk explores the role of Kubernetes, an open-source container orchestration platform, in addressing research data management challenges. I will discuss how we are using a set of publicly available open-source and home-grown tools in the National Renewable Energy Lab (NREL) Data, Analysis, and Visualization (DAV) group to help researchers overcome data-related bottlenecks. The talk will begin by providing an overview of the data challenges faced in research data management, including data storage, processing, and analysis. I will highlight Kubernetes' ability to handle large-scale data by leveraging containerization and distributed computing, including distributed storage. Kubernetes allows researchers to encapsulate data processing infrastructure and workflows into portable containers, enabling reproducibility and ease of deployment. Kubernetes can then schedule and manage the resource allocation of these containers to enable efficient utilization of limited computing resources, leading to more efficient data processing and analysis. I will discuss some limitations of traditional, siloed approaches to dealing with data and emphasize the need for solutions which foster collaboration. I will highlight how we are using Kubernetes at NREL to facilitate data sharing and cooperation among research teams. Kubernetes' flexible architecture enables the deployment of shared computing environments, such as Apache Superset, where researchers can seamlessly access and analyze shared datasets. Providing the ability to have one research team easily consume data generated by another, utilizing Kubernetes' as a central data platform, is one of the major wins we've encountered by adopting the platform. Finally, I will showcase real-world use cases from NREL where we have used Kubernetes to solve some persistent data challenges involving large volumes of sensor and monitoring data. I will discuss the challenges we encountered when creating our cluster and making it available as a production-ready resource. I will also discuss the specific suite of tools, including Postgres and Apache Druid for columnar and timeseries data, and Redpanda Kafka for streaming data we have deployed in our infrastructure, and the process that went into the selection of these tools.

collaborative environment↗

TAMM: Tensor algebra for many-body methods

Tensor algebra operations such as contractions in computational chemistry consume a significant fraction of the computing time on large-scale computing platforms. The widespread use of tensor contractions between large multi-dimensional tensors in describing electronic structure theory has motivated the development of multiple tensor algebra frameworks targeting heterogeneous computing platforms. In this paper, we present Tensor Algebra for Many-body Methods (TAMM), a framework for productive and performance-portable development of scalable computational chemistry methods. TAMM decouples the specification of the computation from the execution of these operations on available high-performance computing systems. With this design choice, the scientific application developers (domain scientists) can focus on the algorithmic requirements using the tensor algebra interface provided by TAMM, whereas high-performance computing developers can direct their attention to various optimizations on the underlying constructs, such as efficient data distribution, optimized scheduling algorithms, and efficient use of intra-node resources (e.g., graphics processing units). The modular structure of TAMM allows it to support different hardware architectures and incorporate new algorithmic advances. We describe the TAMM framework and our approach to the sustainable development of scalable ground- and excited-state electronic structure methods. We present case studies highlighting the ease of use, including the performance and productivity gains compared to other frameworks.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Cooperative Load Scheduling for Multiple Aggregators Using Hierarchical ADMM

Demand response (DR) serves an important role in improving the efficiency and stability of power systems. In recent years, with advances in communication and smart device technologies, many aggregators have emerged to facilitate end customer participation in DR programs. These aggregators, equipped with customized optimal control algorithms, are capable of providing various grid services. Among them is load scheduling during DR events, namely following a load signal provided by the utility company while minimizing overall customer discomfort. However, as the number of aggregators keeps increasing, it becomes challenging for utility companies to conduct load scheduling for multiple aggregators and generate reference signals for each of them. This paper proposes an optimization framework using hierarchical alternating direction method of multipliers (H-ADMM) to optimally generate load following signals for multiple aggregators. Under this framework, utility and multiple aggregators work in a cooperative manner, aiming at minimizing an overall system cost from different levels of the power system hierarchy, while protecting user privacy. A case study has been conducted in a system with multiple aggregators, based on control of HVAC loads. Experimental results validate the effectiveness of the proposed algorithm.

97 MATHEMATICS AND COMPUTING↗

Modular performance prediction for scientific workflows using Machine Learning

Scientific workflows provide an opportunity for declarative computational experiment design in an intuitive and efficient way. A distributed workflow is typically executed on a variety of resources, and it uses a variety of computational algorithms or tools to achieve the desired outcomes. Such a variety imposes additional complexity in scheduling these workflows on large scale computers. As computation becomes more distributed, insights into expected workload that a workflow presents become critical for effective resource allocation. In this paper, we present a modular framework that leverages Machine Learning for creating precise performance predictions of a workflow. The central idea is to partition a workflow in such a way that makes the task of forecasting each atomic unit manageable and gives us a way to combine the individual predictions efficiently. We recognize a combination of an executable and a specific physical resource as a single module. This gives us a handle to characterize workload and machine power as a single unit of prediction. Overall, our modular technique of creating atomic modules and deployment of longest-path approach to estimate workflow performance, allows the framework to adapt to highly complex nested directed acyclic workflows and scale to new scenarios, since it does not make assumptions of underlying workflow structure. We present performance estimation results of independent workflow modules executed on the XSEDE SDSC Comet cluster using various Machine Learning algorithms. The results provide insights into the behavior and effectiveness of different algorithms in the context of scientific workflow performance prediction.

97 MATHEMATICS AND COMPUTING↗

Net Present Value Optimization of a Natural Gas Combined Cycle Plant with CO 2 Capture using a Water-Lean Solvent Considering Transient Electricity Price for Multiple Regions

Global CO 2 emissions are increasing at about a 1.5% rate per year. Fossil fuel-based plants are one of the main contributors to this rise. In the power generation industry, fossil fuel plants are dominant, and many plants are under development. In this study, a natural gas combined cycle (NGCC) power plant with postcombustion capture using a leading water-lean solvent is considered. For optimal design and operating schedule, large-scale dynamic optimization is undertaken for net present value (NPV) optimization. The first principle dynamic model of NGCC is developed, including a model of the highly efficient H-class gas turbines. For computational tractability of the dynamic optimization problem, a reduced-order model is developed by using the Hankel singular value decomposition. A waterlean solvent, N-(2-ethoxyethyl)-3-morpholinopropan-1-amine, is used for carbon capture. A model of the capture system is developed in Aspen Plus, which is used to develop a reduced-order model by using ALAMO, a machine learning software. In addition, a reduced model of the CO 2 compression system with a dehydration unit is also considered. The integrated system is used for NPV optimization by using the Python-based PYOMO platform. The PCC process is analyzed for three configurations-conventional packed bed, rotating packed bed (RPB), and a combination of RPB and direct contact cooler. The NPV optimization is performed for 14 regional markets by considering year-long clustered and continuous locational marginal price data with a 1 h interval. Optimization results show that the PCC can achieve 90% CO 2 capture with a positive NPV for six regions. Sensitivity studies conducted by using the PCC configurations indicate that the process is economically feasible for 9 regions out of 14 regional electricity markets with NPV values in the range of 33−540 $MM.

cabon capture↗

SMC 2021 : Analyzing Resource Utilization and User Behavior on Titan Supercomputer

Resource utilization statistics of submitted jobs on a supercomputer can help us understand how users from various scientific domains use HPC platforms and better design a job scheduler. We explore to generate insight regarding workload distribution and usage pattern domains from job scheduler trace, GPU failure information, and project-specific information collected from Titan supercomputer. Furthermore, we want to know how the scheduler performance varies over time and how the users' scheduling behavior changes following a system failure. These observations have the potential to provide valuable insight, which is helpful to prepare for system failures. These practices will help us develop and apply novel machine learning algorithms in understanding system behavior, requirement, and better scheduling of HPC systems. There are two datasets, RUR and GPU. RUR: This dataset is the job scheduler traces collected from the Titan supercomputerfrom 01/01/2015 to 07/31/2019 (2015.csv - 2019.csv). These were collected usingResource Utilization Report (RUR), a Cray-developed resource-usage data collectionand reporting system. It contains the usage information of its critical resources (CPU,Memory, GPU, and I/O) of each running job on Titan during that period [2]. ProjectAreas: Every job is associated with a project ID. TheProjectAreas.csvdatasetprovides a mapping of the project ID to its domain science. GPU: There have been some hardware-related issues in the GPUs in Titan that caused some GPUs to fail, sometimes irrecoverably during some job runs. This dataset provides information regarding these failures during the execution of the submitted jobs. GPUs on Titan are uniquely identified by a serial number (SN), and they are installed in a location. A GPU can be installed in a location, then removed from that location following a failure, and then re-installed in a different location after fixing the problem. If the failure can't be recovered, the GPU might be removed entirely from Titan. There are two prominent types of failures that resulted in the removal of GPUs from Titan: Double Bit Error (DBE) and Out of the Bus (OTB). The dataset (gc_full.csv) has the following fields: 1. SN : Serial number of a GPU 2. location : The location where it is installed 3. insert : The time when it was inserted into that location 4. remove : The time when it was removed from that location 5. duration : Amount of time the GPU spent in this location 6. out : If the device was taken out entirely w/o a re-installment into a new location. 7. event : If the GPU was taken out entirely, the reason for its removal. To learn more about this dataset, please refer to the git repositoryhttps://github.com/olcf/TitanGPULifeand the related publication [1]. References [1] George Ostrouchov, Don Maxwell, Rizwan A Ashraf, Christian Engelmann, MallikarjunShankar, and James H Rogers. Gpu lifetimes on titan supercomputer: Survival analysisand reliability. InSC20: International Conference for High Performance Computing,Networking, Storage and Analysis, pages 1-14. IEEE, 2020. [2] Feiyi Wang, Sarp Oral, Satyabrata Sen, and Neena Imam. Learning from five-yearresource-utilization data of titan system. In2019 IEEE International Conference onCluster Computing (CLUSTER), pages 1-6. IEEE, 2019.

42 ENGINEERING↗