Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “load balancing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

SIERRA Multimechanics Module: Aria Thermal Theory Manual (V.5.10)

Aria is a Galerkin finite element based program for solving coupled-physics problems described by systems of PDEs and is capable of solving nonlinear, implicit, transient and direct-to-steady state problems in two and three dimensions on parallel architectures. The suite of physics currently supported by Aria includes thermal energy transport, species transport, and electrostatics as well as generalized scalar, vector and tensor transport equations. Additionally, Aria includes support for manufacturing process flows via the incompressible Navier-Stokes equations specialized to a low Reynolds number ($Re$ < 1) regime. Enhanced modeling support of manufacturing processing is made possible through use of either arbitrary Lagrangian-Eulerian (ALE) and level set based free and moving boundary tracking in conjunction with quasi-static nonlinear elastic solid mechanics for mesh control. Coupled physics problems are solved in several ways including fully-coupled Newton’s method with analytic or numerical sensitivities, fully-coupled Newton-Krylov methods and a loosely-coupled nonlinear iteration about subsets of the system that are solved using combinations of the aforementioned methods. Error estimation, uniform and dynamic $h$-adaptivity and dynamic load balancing are some of Aria’s more advanced capabilities.

42 ENGINEERING↗

Implementation and Demonstration of P4 Software for Improving ICS Protocol Visibility and Control [Slides]

No prior enabling funded work applicable to this proposal. Programming Protocol-independent Packet Processors (P4) is an open source, domain-specific programming language for network switching devices. P4 complements traditional Software Defined Networking (SDN) which is primarily concerned with the management of packets (e.g. routing/dropping decisions) rather than how each packet is processed. The introduction of P4 provided new capabilities (e.g. firewall, load balancing, enhanced security) but has primarily been deployed in data centers. This effort investigates ways to expand P4 into other niches such as ICS networks.

97 MATHEMATICS AND COMPUTING↗

Leveraging Computational Storage Devices in Campaign Storage [Slides]

Computational storage provides new ways of accelerating data-intensive applications. In-drive data management schemes matter (O_DIRECT, clustered index). Layer violation: “cheating” one filesystem may be possible; cheating multiple layers of filesystems is hard (FS internal load balancing, fail over, compression, concurrency control). The future directions include block-based acceleration to object-based acceleration.

97 MATHEMATICS AND COMPUTING↗

Deployment of inference as a service at the US CMS Tier-2 data centers

Coprocessors, especially GPUs, will be a vital ingredient of data production workflows at the HL-LHC. At CMS, the GPU-as-a-service approach for production workflows is implemented by the SONIC project (Services for Optimized Network Inference on Coprocessors). SONIC provides a mechanism for outsourcing computationally demanding algorithms, such as neural network inference, to remote servers, where requests from multiple clients are intelligently distributed across multiple GPUs by a load-balancing service. This talk highlights the recent progress in deploying SONIC at selected U.S. CMS Tier-2 data centers. Using realistic CMS Run3 data processing workflows, such as those containing transformer-based algorithms, we demonstrate how SONIC is integrated into the production-like environment to enable accelerated inference offloading. We will present developments from both the client and server sides, including production job and data center configurations for NVIDIA and AMD GPUs. We will also present performance scaling benchmarks and discuss the challenges of operating SONIC in CMS production, such as server discovery, GPU saturation, fallback server logic, etc.

Holzman, Burt↗

Exploring DAOS as a Burst Buffer for a 100 Gbps DAQ Real-Time Streaming System

We present an experimental evaluation of a burst buffer for a real-time DAQ streaming system designed to transmit instrument data to remote data centers. The system is based on EJ-FAT, a load balancing system capable of Nx 100Gbps streams, distributing data from event sources to processing nodes. We explore applying the DAOS system as a burst buffer to serve a number of purposes: improve resiliency, elasticity and add new functions into the processing pipeline. In the evaluation a sender transmits events over a 100Gbps network to a receiver integrated with DAOS to store the reassembled events using DAOS APIs. We evaluate the system for possible bottlenecks and provide end-to-end evaluation with a burst buffer using DAOS storage abstractions. We show that a receiver node can support 38.1 Gbps. This proves the viability of our approach and allows us to extend this work to investigate scale-out properties and new streaming optimizations.

Mei, Xinxin↗

An Efficient Storage-Driven Machine Learning Model for Performance in the Era of Multimodal Scientific Data

Scientific workflows are increasingly relying on machine learning (ML), simulation, and hybrid techniques to predict, understand, and optimize the behavior of complex experiments. High-performance computing has greatly improved researchers’ ability to acquire diverse data modalities in these workflows. Recent studies suggest that the performance of machine learning models can be improved by integrating data from various sources. Unfortunately, these workloads pose unprecedent pressure on the network storage to meet the demands associated with accessing these multimodal data. To mitigate the impact of intensive IO, we propose a solution that utilizes a multi-tier High-Performance Computing (HPC) distributed storage and data processing framework, placing computation where the data resides for better performance. By adopting this project, the scientific community will gain new opportunities to explore multimodal storage-driven possibilities, integrating multiple scientific data sources with advanced streaming frameworks. Additionally, our framework effectively utilizes computing resources and bridges the gaps identified by HPC experts. Our proposed approach tackles scalability and persistence challenges by leveraging native persistency, which has posed difficulties in traditional approaches. Furthermore, we seek to enhance fault-tolerance and load-balance of computations by leveraging real-time streaming in diverse scientific computing environments, thereby propelling advanced scientific computing research into the next generation.

97 MATHEMATICS AND COMPUTING↗

Advancing Grid Resilience through Smart Charge Management: Findings from Maryland’s Pilot

This report presents research findings from a four-year Smart Charge Management (SCM) pilot program conducted by Maryland’s largest electric utilities—Baltimore Gas and Electric (BGE), Potomac Electric Power Company (Pepco), and Delmarva Power & Light (DPL)—to evaluate strategies for optimizing electric vehicle (EV) charging loads and enhancing grid stability. Supported by the U.S. Department of Energy (DOE), Argonne National Laboratory collaborated with all project partners and examined the effectiveness of Time-of-Use (TOU) and Load Balancing (LB) strategies in managing peak demand, deferring costly infrastructure upgrades, and reducing grid constraints at the feeder level.

24 POWER TRANSMISSION AND DISTRIBUTION↗

EJFAT Scientific Perspective

Presented new computing model to the test by deploying the EJFAT system alongside a data-stream processing framework running the production-level CLAS12 event reconstruction application. In this experiment, a continuous stream of CLAS12 Level-1 identified events was processed in real-time using the EJFAT load balancer, distributing the workload across 90 computing nodes located across the U.S. This marks the first-ever large-scale, real-time distributed data stream processing experiment, demonstrating that scientific data-streaming pipelines can efficiently scale across four dimensions, thanks to EJFAT’s advanced hardware and software capabilities.

Gyurjyan, Vardan [Thomas Jefferson National Accele↗

Synchronous Machine Governor Upgrade

Conventional generation sources play a critical role in the stability and reliability of the electrical grid, particularly as we transition towards more renewable energy sources. To understand and accurately emulate their behavior for optimizing grid operations and ensuring seamless integration with renewable technologies, it is essential to better emulate the grid- and plant-level impacts of conventional generation sources, such as natural gas (NG) driven heat recovery steam generators (HRSGs) and combustion turbines (CTs). Therefore, a governor model is developed in a programmable logic controller (PLC) to investigate the performance of the conventional generator under various dynamic operating conditions and to identify the impact on grid stability in a controlled environment. The governor model aims to enable the hardware-in-the-loop (HIL) based emulation of these conventional generation sources using the existing 2 MVA synchronous machine/generator that is driven by a flexible 2.5 MW variable speed drive. This setup will allow us to replicate the dynamic characteristics and response behaviors of NG-driven HRSGs and CTs. The controls for the emulated conventional plants follow the industry standard and are adjustable, ensuring they accurately reflect the operational capabilities and limitations of real-world systems. These controls include load-following capabilities, ramp rates, startup and shutdown sequences, and emissions characteristics. By incorporating these adjustable controls, we aim to capture the nuanced impacts of conventional generation, such as their ability to provide ancillary services like frequency regulation, voltage support, and spinning reserve. In this report, we simulate two types of dynamic operations: grid-connected and islanding. For each dynamic operation, representative starting sequences are tested, including turbine purge, ignition, speed ramping up, generator excitation and synchronizing, and breaker close. The HIL based tests provides insights for field deployment, specifically the high-fidelity governor model provides results to predict the potential stability and reliability risk and suggest possible integration measures (e.g., generation and load balancing, tuning of governor control parameters). Ultimately, this enhanced emulation capability will be integrated into our Advanced Research on Integrated Energy Systems (ARIES), enabling us to conduct comprehensive studies on the interactions between conventional and renewable energy sources. By better understanding these interactions, we can develop strategies to optimize the overall performance and reliability of the grid. This will support the deployment of advanced grid management techniques, such as demand response, grid-forming inverters, and energy storage systems. The main contributions are summarized as follows: (1) This report introduces a PLC-based governor model for gas turbines. This model accurately simulates the dynamic behavior of conventional generation sources under various operational scenarios; (2) The model is integrated with an HIL testbed that includes a 2.5 MW variable speed drive and a 2 MVA synchronous machine. This setup enables realistic, real-time emulation of conventional power plants, particularly NG driven HRSGs and CTs; (3) The developed model is adaptable to various gas turbine configurations and allows for precise control over parameters such as MW ramp rates. This flexibility makes it a valuable tool for future research and industry collaboration; and (4) By incorporating the model into the National Renewable Energy Laboratory's Advanced Research on Integrated Energy Systems, the report lays the groundwork for future studies on interactions between conventional and renewable energy sources, enhancing the ability to develop advanced grid management strategies.

24 POWER TRANSMISSION AND DISTRIBUTION↗

DUNE Rucio Server Scalabiilty Studies

The DUNE collaboration has an ongoing production effort to simulate the full detectors and to analyze the various prototypes that are currently running. Rucio is used to manage the 40PB of files made to date. When 500 or more jobs were sending output to Rucio simultaneously via Rucio upload, we observed timeouts, unhandled exceptions, and Rucio server restarts due to slow performance. In collaboration with the core Rucio team we did a full review of the Rucio upload code and identified several optimizations that can be made. We also have deployed the Ingress load balancer in front of our Rucio servers and added a database connection pooling utility. These changes led to significant improvement both in reliability and scalability, yet we anticipate even better performance will eventually be required. We describe in this paper the initial state of the system, the various debugging processes that were used, and our plans to further improve scalability.

Calcutt, J. [Brookhaven Natl. Lab.]↗

Fast GPU-Based Generation of Large Graph Networks From Degree Distributions

Synthetically generated, large graph networks serve as useful proxies to real-world networks for many graph-based applications. The ability to generate such networks helps overcome several limitations of real-world networks regarding their number, availability, and access. Here, we present the design, implementation, and performance study of a novel network generator that can produce very large graph networks conforming to any desired degree distribution. The generator is designed and implemented for efficient execution on modern graphics processing units (GPUs). Given an array of desired vertex degrees and number of vertices for each desired degree, our algorithm generates the edges of a random graph that satisfies the input degree distribution. Multiple runtime variants are implemented and tested: 1) a uniform static work assignment using a fixed thread launch scheme, 2) a load-balanced static work assignment also with fixed thread launch but with cost-aware task-to-thread mapping, and 3) a dynamic scheme with multiple GPU kernels asynchronously launched from the CPU. The generation is tested on a range of popular networks such as Twitter and Facebook, representing different scales and skews in degree distributions. Results show that, using our algorithm on a single modern GPU (NVIDIA Volta V100), it is possible to generate large-scale graph networks at rates exceeding 50 billion edges per second for a 69 billion-edge network. GPU profiling confirms high utilization and low branching divergence of our implementation from small to large network sizes. For networks with scattered distributions, we provide a coarsening method that further increases the GPU-based generation speed by up to a factor of 4 on tested input networks with over 45 billion edges.

97 MATHEMATICS AND COMPUTING↗

Editorial: Advanced water splitting technologies development: Best practices and protocols

As the level of deployment and utilization of renewable energy sources, including wind and solar, continues to rise, large-scale, long-term energy storage technologies that could accommodate weekly and seasonal energy fluctuations will play a significant role in the overall deployment of renewable energies in the future. Harnessing and storing renewable energy resources via electrochemical, photoelectrochemical, or thermochemical processes by converting renewable energy into sustainable (energy storage) fuels have the potential to meet the long-term, terawatt scale energy storage challenge. Renewable hydrogen production is the cornerstone for sustainable fuel production and deep decarbonization of multiple sectors in our society. Cost-competitive clean hydrogen provides value to applications, such as 1) in the transportation sector for fuel cell vehicles, 2) in the electric grid sector for system stability and load balancing, and 3) in the industrial sector with metal refineries, cement production, and biomass upgrading (carbon-free fertilizer production). In addition, coupling clean renewable hydrogen with the carbon and nitrogen cycles enables known and well-established thermal-chemical processes to generate renewable hydrocarbon fuels and ammonia. The Advanced Water Splitting Technologies (AWST): low temperature electrolysis (LTE), high temperature electrolysis (HTE), photoelectrochemical (PEC) and solar thermo-chemical hydrogen (STCH) provide four unique and parallel approaches to produce low cost, low greenhouse gas (GHG) emission hydrogen at scale (Figure 1). Cost competitive clean hydrogen production using these four technologies is a current high priority focus for governments and industry. In June of 2022, the U.S. Department of Energy (DOE) launched the first in a series of Earthshot Initiatives. The Hydrogen Shot, “1 1 1” aims to reduce the cost of clean hydrogen by more than 80% to one dollar per one kilogram in 1 decade ($\$$1/kg H 2 ). The European Green Deal and the International Energy Agency (IEA) have implemented a strong focus on green hydrogen production for a clean and secure energy future.

benchmarking, low temperature electrolysis↗

An Orthogonal Recursive Bisection (ORB) Based Time Advancement Algorithm for CFD-DEM Solvers

The time integration of the granular phase in coupled computational fluid dynamics (CFD) – discrete element method (DEM) simulations presents a unique computational challenge brought about by the large variations in particle collisional time scales. Particles in the dilute regions of the computational domain can be advanced with large time steps while dense regions require much smaller time increments. However, the time step size in most solvers is globally set as the limit for accuracy and stability imposed by the collisions and is typically orders of magnitude less than that required away from collisions. This work addresses this precise issue and provides a strategy to avoid the use of a global conservative small time step size for the entire set of particles.A novel time stepping algorithm for CFD-DEM solvers using a partitioning approach using orthogonal recursive bisection (ORB) that allows for variable time steps among particles is described and its computational performance is compared against baseline explicit methods, typically used in several CFD-DEM solvers. ORB has advantages of being relatively quick and easy to update incrementally and has the required heuristic behavior (i.e., it will split the region in half with a cluster on each side) when groups of particles are well separated (clustered). The algorithm presented in this work uses a local time stepping approach to resolve collisional time scales for subsets of particles that are present at the leaves of the ORB, thereby resulting in substantial reduction of computational cost. The parallel implementation of this method where a ``knapsack” algorithm is used in tandem with ORB for effective load-balancing is also presented, where a best possible partitioning is obtained based on number of particles and local time-stepping costs. The algorithm is tested against benchmark problems with varying particle distributions that include fluidized bed and riser flow scenarios. Preliminary results indicate that the approach is 2-3X faster than traditional explicit methods for problems that involve both dense and dilute regions, while maintaining the same level of accuracy.

adaptive timestepping↗

Systems and methods for detecting and mitigating cyber attacks on power systems comprising distributed energy resources

Extensive deployment of interoperable distributed energy resources (DER) on power systems is increasing the power system cybersecurity attack surface. National and jurisdictional interconnection standards require DER to include a range of autonomous and commanded grid-support functions which can drastically influence power quality, voltage, and the generation-load balance. Investigations of the impact to the power system in scenarios where communications and operations of DER are controlled by an adversary show that each grid-support function exposes the power system to distinct types and magnitudes of risk. The invention provides methods for minimizing the risks to distribution and transmission systems using an engineered control system which detects and mitigates unsafe control commands.

97 MATHEMATICS AND COMPUTING↗

HEPnOS: a Specialized Data Service for High Energy Physics Analysis

In this paper, we present HEPnOS, a distributed data service for managing data produced by high-energy physics (HEP) experiments. Using HEPnOS, HEP applications can use HPC resources more effciently than traditional fle-based applications. The fle-based model leads to a rigid, chunk-based allocation of computational resources and limits the number of cores that can be used concurrently by an HEP application. The fundamental problem is that organizing domain-specifc data into fles inadvertently introduces a single, artifcial, confated tuning parameter that puts key optimization goals into confict: larger fle sizes reduce metadata overhead and thus improve I/O effciency, but smaller fle sizes provide more opportunity for workfow parallelism and load balancing. In this work, we introduce a domain-specifc data service that decouples that constraint so that data can be accessed and processed in its natural granularity while still maintaining I/O effciency. By removing the constraints introduced by fle handling we are able to obtain better scaling and make effcient use of more cores for processing a fxed-sized data sample. We demonstrate the improved scalability by using an application developed in the fle-based paradigm and comparing it to a version modifed to use HEPnOS.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Unlocking the price

To help balance load and power generation that ensue from the electrification of transportation and the increased connection of variable power sources to the grid, ABB has developed an RL-based dynamic price model for EV charging with the required flexibility to respond to changing grid conditions.

Suryanarayana, Harish↗

Systems and methods for detecting and mitigating cyber attacks on power systems comprising distributed energy resources

Extensive deployment of interoperable distributed energy resources (DER) on power systems is increasing the power system cybersecurity attack surface. National and jurisdictional interconnection standards require DER to include a range of autonomous and commanded grid-support functions which can drastically influence power quality, voltage, and the generation-load balance. Investigations of the impact to the power system in scenarios where communications and operations of DER are controlled by an adversary show that each grid-support function exposes the power system to distinct types and magnitudes of risk. The invention provides methods for minimizing the risks to distribution and transmission systems using an engineered control system which detects and mitigates unsafe control commands.

Johnson, Jay Tillay↗