Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “cluster scheduling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Renewable hydrogen and ammonia for combined heat and power systems in remote locations: Optimal design and scheduling

Abstract Using hydrogen (H ) and ammonia (NH ) for renewable energy storage has the potential to enable economical power and heat supply with high renewable penetrations, especially in remote locations which are characterized by high energy costs. In this work we assess the economic competitiveness of renewable combined heat and power (CHP) systems in Mahaka HI, Nantucket MA, and Northwest Arctic Borough (NWAB) AK by optimally designing these systems for scenarios in which power and heat can be purchased over a range of historical energy prices as well as when 100% renewable supply is required. We use a combined optimal design and scheduling model which minimizes annualized net present cost by determining optimal technology selection and size simultaneously with optimal schedules for each period of a system operating horizon aggregated from full year hourly resolution data via a consecutive temporal clustering algorithm. We find that renewable generation meets at least 85% of power demands and 75% of heat demands under the lowest energy prices investigated. Higher conventional energy prices lead to increased renewable penetration which is facilitated by renewable NH as a seasonal energy storage medium, as are 100% renewable CHP systems. NH is used for power generation with heat cogeneration in all three locations, as well as directly for heating in NWAB. On an annual cost basis, NH ‐enabled 100% renewable CHP is only 3% more expensive in Mahaka and NWAB than systems which can purchase energy at the lowest prices, while it is 15% more expensive in Nantucket.

Palys, Matthew J.↗

Condition-Based Maintenance of a Circulating Water System of a Canadian Nuclear Power Plant using Machine Learning and Statistical Tools

Canada Deuterium Uranium pressurized-heavy-water reactors (PHWR) are a type of nuclear power plant that generate clean and reliable energy. The scope of this work is to automate data analysis methodologies to inform a condition-based maintenance strategy of a circulating water system (CWS) of a PHWR. The multiunit CWS provides a continuous supply of water to cool steam condensers, even during transient scenarios, thereby improving the thermal efficiency. This work aims to develop a machine learning (ML) based approach to detect anomalies in heterogeneous data of a CWS in a PHWR to help inform a predictive maintenance strategy. The heterogeneous data include textual and numeric time series data for a PHWR. Natural-language-processing (NLP)-based models are used to analyze textual data contained in work orders and operator logs and an event-timeseries correlation detection method is applied to assist anomalies diagnoses for CWS. An ML model Robust Linear Model (RLM) is also used to remove the seasonal variations in the system variable distributions based on distributions of environmental variables. A machine learning model, Density-Based Spatial Clustering of Applications with Noise (DBSCAN), trained on both original data and data without any seasonal variations will then be used to detect if an anomaly exists. Thus, by moving to an automated methodology to detect, classify, and forecast anomalies, the maintenance strategy would be based on component condition instead of a time-based schedule.

97 - MATHEMATICS AND COMPUTING↗

Efficient prediction of concentrating solar power plant productivity using data clustering

Concentrating solar power (CSP) plants convert solar energy to electricity and can be deployed with a thermal storage capability to shift electricity generation from time periods with available solar resource to those with high electricity demand or electricity price. Rigorous optimization of plant design and operational strategies can improve the market-competitiveness and commercial viability; however, such optimization may require hundreds of annual performance simulations, each of which can be computationally expensive when including considerations such as optimization of dispatch scheduling, sub-hourly time resolution, and stochastic effects due to uncertain weather or electricity price forecasts. This paper proposes a methodology to reduce the computational burden associated with simulation of electricity yield and revenue for CSP plants over a single- or multi-year period. Data-clustering techniques are employed to select a small number of limited-duration time blocks for simulation that, when appropriately weighted, can reproduce generation and revenue over a single year or within each year of a multi-year period. After selection of appropriate data features and weighting factors defining similarity between time-series profiles, the methodology captured annual revenue within 2.3%, 1.7%, or 1.2% using simulation of 10, 30, or 50 three-day exemplar time blocks, respectively, for each of three single-year location/weather/market scenarios and five plant configurations ranging from low to high solar multiple and storage capacity. When applied to multi-year datasets, the proposed methodology can capture inter-year variability that is unavailable from typical meteorological year (TMY) datasets while simultaneously requiring simulation of less than a single year of data.

14 SOLAR ENERGY↗

Profiles of upcoming HPC Applications and their Impact on Reservation Strategies

With the expected convergence between HPC, BigData and AI, new applications with different profiles are coming to HPC infrastructures. Here, we aim at better understanding the features and needs of these applications in order to be able to run them efficiently on HPC platforms. The approach followed is bottom-up: we study thoroughly an emerging application from the neuroscience community (SLANT) to understand its behavior. Based on these observations, we derive a generic, yet simple, application model (namely, a linear sequence of stochastic jobs). We expect this model to be representative for a large set of upcoming applications that require the computational power of HPC clusters without fitting the typical behavior of large-scale traditional applications. In a second step, we show how one can manipulate this generic model in a scheduling framework. Specifically we consider the problem of making reservations (both time and memory) for an execution on an HPC platform. We derive solutions using the model of the first step of this work. We experimentally show the robustness of the model, even with very few data or with another application, to generate the model, and provide performance gains with regards to standard and more recent approaches used in the neuroscience community.

97 MATHEMATICS AND COMPUTING↗

EJFAT: Towards Intelligent Compute Destination Load Balancing

To handle increased data flow, Jefferson Lab (JLab) is partnering with ESnet for development of an AI/ML directed compute work Load Balancer (LB) of UDP streamed data. The LB is FPGA based featuring dynamically configurable, low latency and high throughput destination address switching. The LB provides integration of edge and core computing to support JLab experimental programs, the Electron-Ion Collider, as well as data centers of the future. In the ESnet/JLab FPGA Accelerated Transport (EJFAT) initiative, the function of the LB Data Plane (DP) is to redirect data streams to selectable (but unknown to sender) destination hosts based on current worload and within that host to destination ports as a function of sub- stream id. This effects hierarchical scaling, first across compute machines for processing over a series of events and second, across ports so different data source sub-streams may be assigned to different processors for further parallelization. The LB Control Plane (CP) programs the DP using compute farm telemetry to direct and balance workloads across a compute cluster as the operating conditions require. While Proportional/Integrative/Derivative (PID) controllers are often seen in similar applications, here we investigate the feasibility of a Reinforcement Learning (RL) based schedule manager running in the CP to provide dynamic updates to the DP scheduling policy.

Lawrence, David↗

New Fracture Diagnostic Tool for Unconventionals: High-Resolution Distributed Strain Sensing via Rayleigh Frequency Shift during Production in Hydraulic Fracture Test 2

Fiber Optic monitoring in unconventional reservoirs has proven to be an invaluable diagnostic tool for assessing both near-wellbore stimulation effectiveness and to help describe the far-field frac geometries created by hydraulic fracture stimulation. Unfortunately, gaining any detailed qualitative and quantitative understanding of the near-wellbore frac geometry or cluster/stage productivity during production via Fiber Optic (FO) has proven to be more difficult, particularly in wells producing liquids. A new FO diagnostic method, Distributed Strain Sensing based on Rayleigh Frequency Shift (DSS-RFS), first demonstrated for oil and gas applications in the Hydraulic Test Site 2 (HFTS2) provides new insights about the characteristics of near-wellbore-region (NWR) during production. DSS-RFS is different from other FO strain measurements because it relies on accurate measurement of frequency shifts of Rayleigh backscattered spectrum obtained by scanning the fiber with a coherent optical time-domain reflectometer with a range of laser frequencies using a tunable-wavelength laser system. Changes in strain are measured with an extremely high spatial resolution of 20 cm and with high signal-to-noise ratios over long distances. In HFTS2, strain changes for the entire wellbore have been measured twice during scheduled shut-in and reopening operations (February 2020 and September 2020). After removing temperature effects, consistent strain changes have been observed at the location of most perforation clusters. These are caused by near wellbore fracture aperture changes due to pressure increases during shut-in within the near-wellbore fracture network. The strain-change patterns from the DSS-RFS during shut-in correlate very well with the location of clusters and allow for the definition of extending intervals with positive strain signals at each cluster and slightly compressing intervals with negative strain signals between the clusters and in the non stimulated intervals. The locations of the measured positive strain peaks also show good correspondence to DAS acoustic intensity measurements acquired during the stimulation. The geometry and magnitude of the strain changes differ significantly between the two tested completion designs in the same well. During shut in and reopening each cluster exhibit its own strain-change / pressure path. In addition, the September 2020 dataset also revealed the existence of small but measurable strain changes as consequence of pressure decline during production. These strain changes also correlate well with the presence of producing clusters, but the strain-rate signals are opposite to that obtained during shut-in and reopening operations. Although Downloaded from http://onepetro.org/URTECONF/proceedings-pdf/21URTC/2-21URTC/D021S031R002/2477551/urtec-2021-5408-ms.pdf/1 by Carol Worster on 28 February 2022 URTeC 5408 2 we are still in the early stages of exploring the potential of this novel FO technique, we believe that the highly detailed information contained in the measurement of strain changes using DSS-RFS during production can significantly improve our understanding of near-wellbore hydraulic fracture characteristics and the relationships between stimulation and production from unconventional oil and gas wells.

04 OIL SHALES AND TAR SANDS↗

Investigation into the transport properties of planetary interiors through inelastic X-ray scattering experiments and quantum molecular dynamics (Final Technical Report)

This was a two-year research project for the period 08/15/2018 - 08/14/2020 conducted at the University of Nevada, Reno using Pronghorn, a newly built high-performance cluster located at the University. The short-term project goal was to provide computational support for the experimental campaigns at LCLS through data analysis of previous and upcoming scheduled experiments, along with performing state-of-the-art atomistic simulations of dense plasmas. The longer-term goal was to create a new computational high energy density physics group located at the University of Nevada, Reno. Guided by results from past and future high-resolution scattering experiments, we aimed to research and develop non-equilibrium and non-adiabatic atomistic simulations of dense plasmas that went beyond the Born-Oppenheimer approximation

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Using hydrogen and ammonia for renewable energy storage: A geographically comprehensive techno-economic study

Hydrogen and, more recently, ammonia have received worldwide attention as energy storage media. In this work we investigate the economics of using each of these chemicals as well as the two in combination for islanded renewable energy supply systems in 15 American cities representing different climate regions throughout the country. We use an optimal combined capacity planning and scheduling model which minimizes the levelized cost of energy (LCOE) by determining optimal unit selection and size along with unit commitments, production rates, and storage inventories for each period of system operation. These periods are aggregated from full year hourly resolution data via a consecutive temporal clustering algorithm. Ammonia is generally more economical than hydrogen as a single method of energy storage. Additionally, systems which use both hydrogen and ammonia outperform those which use only one storage option and have LCOE between $\$ 0.17$/kWh and $\$ 0.28$/kWh, including full investment in renewable generation infrastructure.

25 ENERGY STORAGE↗

Fracture Network Prediction Using Physics-based Machine Learning Algorithms

In recent years, systematic CO2 injection into geological reservoirs across the U.S. has gained traction as a strategy to mitigate greenhouse gas emissions. This approach necessitates precise monitoring to ensure secure containment, minimize risks, and optimize storage management. Our study leverages machine learning (ML) techniques to advance the understanding of CO2 injection processes, focusing on the Illinois Basin. Over a three-year injection period, we analyzed microseismic data, identifying 19 temporal intervals with significant bottom-hole pressure changes. By partitioning microseismic events into these intervals and estimating b-values, we revealed over 100 clusters of events related to fracture initiation or reactivation. Advanced spatial analysis highlighted horizontally-oriented fractures along the NNW-SSE axis. This quantification of fracture networks informs dynamic injection scheduling, work-over strategies, and risk assessments, enhancing carbon capture, utilization, and storage (CCUS) operations. Additionally, our methodology offers valuable insights for oil and gas operations and geothermal development, supporting fracture-based monitoring and risk mitigation.

Kumar, Abhash↗

Electric vehicle supply equipment location and capacity allocation for fixed-route networks

Electric vehicle (EV) supply equipment location and allocation (EVSELCA) problems for freight vehicles are becoming more important because of the trending electrification shift. Some previous works address EV charger location and vehicle routing problems simultaneously by generating vehicle routes from scratch. Although such routes can be efficient, introducing new routes may violate practical constraints, such as drive schedules, and satisfying electrification requirements can require dramatically altering existing routes. To address the challenges in the prevailing adoption scheme, we approach the problem from a fixed -route perspective. We develop a mixed -integer linear program, a clustering approach, and a metaheuristic solution method using a genetic algorithm (GA) to solve the EVSELCA problem. The clustering approach simplifies the problem by grouping customers into clusters, while the GA generates solutions that are shown to be nearly optimal for small problem cases. A case study examines how charger costs, energy costs, the value of time (VOT), and battery capacity impact the cost of the EVSELCA. Charger equipment costs were found to be the most significant component in the objective function, leading to a substantial reduction in cost when decreased. VOT costs exhibited a significant decrease with rising energy costs. Further, an increase in VOT resulted in a notable rise in the number of fast chargers. Longer EV ranges decrease total costs up to a certain point, beyond which the decrease in total costs is negligible.

33 ADVANCED PROPULSION SYSTEMS↗

Elastic distributed training with fast convergence and efficient resource utilization

Distributed learning is now routinely conducted on cloud as well as dedicated clusters. Training with elastic resources brings new challenges and design choices. Prior studies focus on runtime performance and assume a static algorithmic behavior. In this work, by analyzing the impact of of resource scaling on convergence, we introduce schedules for synchronous stochastic gradient descent that proactively adapt the number of learners to reduce training time and improve convergence. Our approach no longer assumes a constant number of processors throughout training. In our experiment, distributed stochastic gradient descent with dynamic schedules and reduction momentum achieves better convergence and significant speedups over prior static ones. Numerous distributed training jobs running on cloud may benefit from our approach.

Cong, Guojing↗

AI4IO: A suite of AI-based tools for IO-aware scheduling

Traditional workload managers do not have the capacity to consider how IO contention can increase job runtime and even cause entire resource allocations to be wasted. Whether from bursts of IO demand or parallel file systems (PFS) performance degradation, IO contention must be identified and addressed to ensure maximum performance. In this paper, we present AI4IO (AI for IO), a suite of tools using AI methods to prevent and mitigate performance losses due to IO contention. AI4IO enables existing workload managers to become IO-aware. Currently, AI4IO consists of two tools: PRIONN and CanarIO. PRIONN predicts IO contention and empowers schedulers to prevent it. CanarIO mitigates the impact of IO contention when it does occur. We measure the effectiveness of AI4IO when integrated into Flux, a next-generation scheduler, for both small- and large-scale IO-intensive job workloads. Our results show that integrating AI4IO into Flux improves the workload makespan up to 6.4%, which can account for more than 18,000 node-h of saved resources per week on a production cluster in our large-scale workload.

Wyatt, II, Michael R.↗

Shared Use Travel Behavior for Improving Rural Mobility: Insights from Greene County, Pennsylvania

Rural communities are considered disadvantaged communities as they suffer from a lack of transport options. Thus, rural regionsprovide less accessibility for commuters to reach their destination as opposed to urban regions. However, the issues of transport disadvantageand shared use mobility in rural areas within the United States (US) have not been well investigated. Furthermore, transport disadvantagediffers between communities and regions across the globe; thus, there is a need to study the behavioral choices of rural commuters within theUS context. This study contributes by analyzing the behavioral choices of rural communities within the US through a case study site ofWaynesburg, Pennsylvania, for adopting a shared use shuttle service. K-means clusters showed that trips from the survey data were a goodrepresentation of real trips from Ecolane. Furthermore, random parameter-based binary logit models were calibrated using data collected fromstudents, faculty, and residents in Waynesburg, Greene County, to study the behavioral choices of commuters. The findings for the faculty andstudents group revealed that prior experience with shared services increases the likelihood of using a shared shuttle. An important personalcharacteristic of inconvenience showed a higher propensity toward using existing modes as opposed to a shared shuttle. Such commutersvalue personal vehicles as more convenient as they have childcare responsibilities and varying schedules for work that require them to moveback and forth across locations, thus making a shared shuttle less attractive for them. The socioeconomic factors of age and gender show ahigher propensity for using shared shuttles. Furthermore, the findings from this study could be helpful for agencies in improving rural mobility andconsidering such shared mobility services for rural communities

42 ENGINEERING↗

Machine Learning–Based Condition Monitoring of a Circulating Water System of a Canadian Nuclear Plant

With the need to maintain long-term reliable energy using nuclear power plants, there is an underlying demand to ensure that the maintenance of plant components and systems is also done in an efficient and cost-effective manner. One way to achieve this is by moving from time-based maintenance to condition-based maintenance. The research presented in this paper focuses on applying statistical and machine-learning-based methods to capture anomalies within data for fault detection to further develop into condition monitoring. This paper focuses on system data for a circulating water system (CWS) of a pressurized heavy-water reactor for detecting anomalies. The different methodologies used for detecting and capturing anomalies in the CWS data are matrix profile, density-based spatial clustering of applications with noise (DBSCAN), and support vector machines (SVMs). Matrix profile and DBSCAN are used to distinguish between normal data and anomalous data. This paper presents a hybrid method using DBSCAN and SVM when a portion of the data is used for DBSCAN to generate clusters. This portion of data is then used to train the SVM along with the clusters generated by DBSCAN as output. SVM is then tested on unseen data as a predictive tool, which can work in real time to categorize data points as either normal or anomalous. This paper presents results that show the high accuracies of DBSCAN and SVM in capturing anomalies within the data for a CWS for fault detection. Thus, the maintenance plan would be focused on component condition rather than a time-based schedule by switching to an automated system to identify and predict faults within a CWS.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Scalable All-pairs Shortest Paths for Huge Graphs on Multi-GPU Clusters

We present an optimized Floyd-Warshall (Floyd-Warshall) algorithm that computes the All-pairs shortest path (APSP) for GPU accelerated clusters. The Floyd-Warshall algorithm due to its structural similarities to matrix-multiplication is well suited for highly parallel GPU architectures. To achieve high parallel efficiency, we address two key algorithmic challenges: reducing high communication overhead and addressing limited GPU memory. To reduce high communication costs, we redesign the parallel (a) to expose more parallelism, (b) aggressively overlap communication and computation with pipelined and asynchronous scheduling of operations, and (c) tailored MPI-collective. To cope with limited GPU memory, we employ an offload model, where the data resides on the host and is transferred to GPU on-demand. The proposed optimizations are supported with detailed performance models for tuning. Our optimized parallel Floyd-Warshall implementation is up to 5x faster than a strong baseline and achieves 8.1 PetaFLOPS/sec on 256~nodes of the Summit supercomputer at Oak Ridge National Laboratory. This performance represents 70% of the theoretical peak and 80% parallel efficiency. The offload algorithm can handle 2.5x larger graphs with a 20% increase in overall running time.

Sao, Piyush↗

Machine Learning Assisted HPC Workload Trace Generation for Leadership Scale Storage Systems

Monitoring and analyzing a wide range of I/O activities in an HPC cluster is important in maintaining mission-critical performance in a large-scale, multi-user, parallel storage system. Center-wide I/O traces can provide high-level information and fine-grained activities per application or per user running in the system. Studying such large-scale traces can provide helpful insights into the system. It can be used to develop predictive methods for making predictive decisions, adjusting scheduling policies, or providing decisions for the design of next-generation systems. However, sharing real-world I/O traces to expedite such research efforts leaves a few concerns; i) the cost of sharing the large traces is expensive due to this large size, and ii) privacy concern is an issue.We address such issues by building an end-to-end machine learn- ing (ML) workflow that can generate I/O traces for large-scale HPC applications. We leverage ML based feature selection and gener- ative models for I/O trace generation. The generative models are trained on I/O traces collected by the darshan I/O characterization tool over a period of one year. We present a two-step generation process consisting of two deep-learning models, called the feature generator and the trace generator. The combination of two-step generative models provides robustness by reducing the bias of the model and accounting for the stochastic nature of the I/O traces across different runs of an application. We evaluate the performance of the generative models and show that the two-step model can generate time-series I/O traces with less than 20% root mean square error.

Paul, Arnab↗

AmeriFlux FLUXNET-1F US-RC1 Cook Agronomy Farm - No Till

This is the AmeriFlux Management Project (AMP) created FLUXNET-1F version of the carbon flux data for the site US-RC1 Cook Agronomy Farm - No Till. This is the FLUXNET version of the carbon flux data for the site US-RC1 Cook Agronomy Farm - No Till produced by applying the standard ONEFlux (1F) software. Site Description - RC1 operated from 2013-2016 at the R.J. Cook Agronomy Farm, as part of a cluster of 5 towers (RC1 to RC5) operated for the Regional Approaches to Climate Change (REACCH) USDA-supported research project. The tower predates the Longterm Agroecosystem Research (LTAR) site common experiment, which was established in nearby fields at the Cook Agronomy Farm in 2017. Cook Agronomy Farm is in the high precipitation agroecological zone of the Columbia Plateau’s dryland cropping region. Wheat-based crop rotations are grown on an annual planting schedule. RC1 was in no-till management since 1998, and was contrasted with RC2, which had conventional, reduced-tillage management. RC1 captured the same tillage practices as the US-CF1 site established in 2017 as part of LTAR common experiment. However, the towers have distinct footprints, aspects, and soil series composition.

Chi, Jinshu [The Hong Kong University of Science a↗

Computing the Properties of Matter with Leadership Computing Resources (Closeout Report for DE-SC0018121)

In order to add more capabilities to Halide, we have designed a new framework called Tiramisu and integrated this framework into Halide. Since Tiramisu enables Halide to target heterogeneous architectures, our development efforts have been refocused on Tiramisu. Most high-performance computer systems today are complex and increasingly heterogeneous; they may have CPUs, GPUs and FPGAs. Achieving best performance requires taking full advantage of all these different architectures. To address this issue, we have designed Tiramisu, an optimization framework that enables Halide (and other DSLs) to target heterogeneous architectures. Tiramisu is an optimization framework that takes as input a high level, architecture-independent representation of code and a set of scheduling and data mapping commands that guide code transformation. The input can either be generated by a domain-specific language (DSL) compiler such as Halide or directly written by a programmer. Tiramisu then applies the user-specified code and data-layout transformations and generates an architecture-specific, low-level intermediate representation (IR) that takes advantage of modern architectural features such as multicore parallelism, non-uniform memory (NUMA) hierarchies, clusters, and accelerators like GPUs and FPGAs. We integrated Tiramisu within Halide and implemented a representative set of benchmarks to evaluate this integration. Tiramisu is now open source and is available for public use (http://tiramisu-compiler.org/). A paper about Tiramisu was published, it shows that Tiramisu extends Halide with many new capabilities and that Tiramisu can generate efficient code for multicores, GPUs, FPGAs and distributed heterogeneous systems. The performance of code generated by the Tiramisu backends matches or exceeds hand optimized reference implementations. For example, the multicore backend matches the highly optimized Intel MKL library on many kernels and shows speedups reaching 4x over the original Halide. In addition to making Tiramisu more robust, we have used Tiramisu to implement a set of representative tensor operation for constructing baryon building blocks required for multi baryon contractions in LQCD. In order to implement this code, we needed to generalize Tiramisu in two ways: first we needed to support indirect array accesses, and second, we needed to add support for complex numbers to Tiramisu. The code generated by Tiramisu is 6x faster than the reference code. Our efforts towards an MPI based multi-node version of tiramisu have matured and the resulting code scales well on multiple nodes (tests up to 512 KNL nodes have been undertaken).

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗