Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “load balancing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

OCTOKV: An Agile Network-Based Key-Value Storage System with Robust Load Orchestration

In this paper, we propose OctoKV, an innovative network-based key-value storage system. OctoKV addresses the repetitive address translation overhead associated with traditional key-value stores running on file systems on the client side. To mitigate this overhead, we implemented the key-value store on the server side using NVMe-oF and a user-level NVMe driver. In particular, we employed fine-grained resource monitoring and load balancing based on heuristics to optimize I/O performance. OctoKV is deployed on a Linux cluster with Intel SPDK. The extensive evaluation shows that OctoKV achieves lower I/O response times in comparison to traditional approaches where key-value stores run on the client side. Also, the proposed load balancing strategies efficiently enhance I/O response times by equally distributing the workload from overloaded cores to other cores.

Khan, Awais↗

A Sparse Distributed Gigascale Resolution Material Point Method

In this paper, we present a four-layer distributed simulation system and its adaptation to the Material Point Method (MPM). The system is built upon a performance portable C++ programming model targeting major High-Performance-Computing (HPC) platforms. A key ingredient of our system is a hierarchical block-tile-cell sparse grid data structure that is distributable to an arbitrary number of Message Passing Interface (MPI) ranks. We additionally propose strategies for efficient dynamic load balance optimization to maximize the efficiency of MPI tasks. Our simulation pipeline can easily switch among backend programming models, including OpenMP and CUDA, and can be effortlessly dispatched onto supercomputers and the cloud. Finally, we construct benchmark experiments and ablation studies on supercomputers and consumer workstations in a local network to evaluate the scalability and load balancing criteria. We demonstrate massively parallel, highly scalable, and gigascale resolution MPM simulations of up to 1.01 billion particles for less than 323.25 seconds per frame with 8 OpenSSH-connected workstations.

97 MATHEMATICS AND COMPUTING↗

Implementation of a fully-balanced periodic tridiagonal solver on a parallel distributed memory architecture

While parallel computers offer significant computational performance, it is generally necessary to evaluate several programming strategies. Two programming strategies for a fairly common problem - a periodic tridiagonal solver - are developed and evaluated. Simple model calculations as well as timing results are presented to evaluate the various strategies. The particular tridiagonal solver evaluated is used in many computational fluid dynamic simulation codes. The feature that makes this algorithm unique is that these simulation codes usually require simultaneous solutions for multiple right-hand-sides (RHS) of the system of equations. Each RHS solutions is independent and thus can be computed in parallel. Thus a Gaussian elimination type algorithm can be used in a parallel computation and the more complicated approaches such as cyclic reduction are not required. The two strategies are a transpose strategy and a distributed solver strategy. For the transpose strategy, the data is moved so that a subset of all the RHS problems is solved on each of the several processors. This usually requires significant data movement between processor memories across a network. The second strategy attempts to have the algorithm allow the data across processor boundaries in a chained manner. This usually requires significantly less data movement. An approach to accomplish this second strategy in a near-perfect load-balanced manner is developed. In addition, an algorithm will be shown to directly transform a sequential Gaussian elimination type algorithm into the parallel chained, load-balanced algorithm.

Eidson, T. M.↗

Investigation of the applicability of a functional programming model to fault-tolerant parallel processing for knowledge-based systems

In a fault-tolerant parallel computer, a functional programming model can facilitate distributed checkpointing, error recovery, load balancing, and graceful degradation. Such a model has been implemented on the Draper Fault-Tolerant Parallel Processor (FTPP). When used in conjunction with the FTPP's fault detection and masking capabilities, this implementation results in a graceful degradation of system performance after faults. Three graceful degradation algorithms have been implemented and are presented. A user interface has been implemented which requires minimal cognitive overhead by the application programmer, masking such complexities as the system's redundancy, distributed nature, variable complement of processing resources, load balancing, fault occurrence and recovery. This user interface is described and its use demonstrated. The applicability of the functional programming style to the Activation Framework, a paradigm for intelligent systems, is then briefly described.

Harper, Richard↗

Parallel DSMC Solution of Three-Dimensional Flow Over a Finite Flat Plate

This paper describes a parallel implementation of the direct simulation Monte Carlo (DSMC) method. Runtime library support is used for scheduling and execution of communication between nodes, and domain decomposition is performed dynamically to maintain a good load balance. Performance tests are conducted using the code to evaluate various remapping and remapping-interval policies, and it is shown that a one-dimensional chain-partitioning method works best for the problems considered. The parallel code is then used to simulate the Mach 20 nitrogen flow over a finite-thickness flat plate. It is shown that the parallel algorithm produces results which compare well with experimental data. Moreover, it yields significantly faster execution times than the scalar code, as well as very good load-balance characteristics.

Nance, Robert P.↗

Earth Independent Medical Operations (EIMO)

Inherent in interplanetary space travel are unprecedented challenges that could threaten mission success and negatively impact crew health and performance. Return to definitive care is essentially untenable and resources will be constrained with practically no re-supply capability. Access to ground-based medical expertise will be significantly delayed under nominal conditions with exacerbation during conjunction or prolonged dust storms. Taken together, these challenges necessitate the development of a progressively autonomous medical operational support system to assist the crew medical officer (CMO). While support from ground based medical experts will remain indispensable for pre-mission planning, the approach to management of acute/emergent medical contingencies will require a gradual transition of medical care and decision making from terrestrial to space-based assets, enabling support of astronaut health and performance and reducing overall mission risk. To progressively enable EIMO, a series of meetings were convened with subject matter experts from within NASA, academia and industry to facilitate mapping of the path to support autonomous medical operations. Topics explored in these meetings included the scope of data (storage capacity, usage, transmission rate and bandwidth, computing capacity), CMO training, supply and resource management and task load balance. Recommendations from these meetings will inform the EIMO Concept of Operations and definition of the associated requirements culminating in updates to the NASA 3001 standards. The EIMO project team will work with stakeholders to conceptualize a clinical decision support system (CDSS) to assist the CMO in response to medical contingencies when terrestrial support is delayed or otherwise unavailable. The CDSS will be a system of systems that will utilize data from numerous input vectors. Successful deployment of the CDSS will facilitate medical decision making while decreasing the cognitive load leading to an improvement in task load balancing. Additional benefits of the envisioned CDSS include assistance with inventory management, locating resources, storage/retrieval of medical records, highlighting trends in recorded data, in addition to providing a consult for diagnosis and treatment. More advanced features might include passive monitoring to identify early warning signs of behavioral or medical anomalies to possibly pre-empt onset of conditions that would compromise crew health and performance.

Benjamin Easter↗

Data Structure and Parallel Decomposition Considerations on a Fibonacci Grid

The Fibonacci grid, proposed by Swinbank and Purser (see companion abstract), provides attractive properties for global numerical atmospheric prediction by offering an optimally homogeneous, geometrically regular, and approximately isotropic discretization, with only the polar regions requiring special numerical treatment. It is a mathematical idealization, applied to the sphere, of the multi-spiral patterns often found in botanical structures, such as in pine cones and sunflower heads. Computationally, it is natural to organize the domain, into zones, in each of which the same pair, or triple, of "Fibonacci spirals" dominate. But the further subdivision of such zones into "tiles" of a shape and size suitable for distribution to the processors of a massively parallel computer requires very careful consideration if the subsequent spatial computations along the respective spirals, especially those computations (such as compact differencing schemes) that involve recursion, can be implemented in an efficient "load-balanced "manner without requiring excessive amounts of inter-processor communications. In this paper we show how certain "number theoretic" properties of the Fibonacci sequence (whose numbers prescribe the multiplicity of successive spirals) may be exploited in the decomposition of grid zones into tidy arrangements of triangular grid tiles, each tile possessing one side approximately parallel to the constant-latitude zone boundary. We also describe how the spatially recursive processes may be decomposed across such a tiling, and the directionality of the recursions reversed on alternate grid lines, to ensure a very high degree of load balancing throughout the execution of the computations required for one time step of a global model.

Michalakes, John↗

T-FSM: A Scalable Distributed Task-Based System for Frequent Subgraph Pattern Mining from a Big Graph

Finding frequent subgraph patterns in a big graph is an important problem with many applications such as classifying chemical compounds and building indexes to speed up graph queries. Since this problem is NP-hard, some recent parallel and distributed systems have been developed to accelerate the mining. However, they often have a huge memory cost, very long running time, suboptimal load balancing, poor scale-out capability, and possibly inaccurate results. In this article, we propose an efficient system called T-FSM for parallel mining of frequent subgraph patterns in a big graph. T-FSM supports a new anti-monotonic frequentness measure called Fraction-Score, which is more accurate than the widely used MNI measure. The execution engine of T-FSM supports both intra-machine parallelism and inter-machine parallelism. For intra-machine parallelism, T-FSM adopts a novel task-based execution model to ensure high multithreading concurrency, bounded memory consumption, and effective load balancing. For inter-machine parallelism, T-FSM ensures good scale-out performance with a lightweight pattern rebalancing approach that reduces workload skewness of pattern evaluations among machines. To avoid recomputing the contexts for migrated patterns, we design a novel context cache table to support concurrent and asynchronous requesting and caching of remote context data, which can timely evict and garbage collect used pattern contexts that are no longer needed to keep memory consumption bounded. Extensive experiments show that T-FSM is orders of magnitude faster than existing state-of-the-art parallel systems (more than 10×, 51×, 131×, 55× speedup over ScaleMine, DistGraph, Pangolin and Peregrine, respectively) and distributed systems (more than 42× and 88× over ScaleMine and DistGraph, respectively) for frequent subgraph pattern mining, and it scales out satisfactorily to 512 CPU cores on the Polaris supercomputer at Argonne National Laboratory.

97 MATHEMATICS AND COMPUTING↗

Parallel implementation and evaluation of motion estimation system algorithms on a distributed memory multiprocessor using knowledge based mappings

Several techniques to perform static and dynamic load balancing techniques for vision systems are presented. These techniques are novel in the sense that they capture the computational requirements of a task by examining the data when it is produced. Furthermore, they can be applied to many vision systems because many algorithms in different systems are either the same, or have similar computational characteristics. These techniques are evaluated by applying them on a parallel implementation of the algorithms in a motion estimation system on a hypercube multiprocessor system. The motion estimation system consists of the following steps: (1) extraction of features; (2) stereo match of images in one time instant; (3) time match of images from different time instants; (4) stereo match to compute final unambiguous points; and (5) computation of motion parameters. It is shown that the performance gains when these data decomposition and load balancing techniques are used are significant and the overhead of using these techniques is minimal.

Choudhary, Alok Nidhi↗

A Multi-Level Parallelization Concept for High-Fidelity Multi-Block Solvers

The integration of high-fidelity Computational Fluid Dynamics (CFD) analysis tools with the industrial design process benefits greatly from the robust implementations that are transportable across a wide range of computer architectures. In the present work, a hybrid domain-decomposition and parallelization concept was developed and implemented into the widely-used NASA multi-block Computational Fluid Dynamics (CFD) packages implemented in ENSAERO and OVERFLOW. The new parallel solver concept, PENS (Parallel Euler Navier-Stokes Solver), employs both fine and coarse granularity in data partitioning as well as data coalescing to obtain the desired load-balance characteristics on the available computer platforms. This multi-level parallelism implementation itself introduces no changes to the numerical results, hence the original fidelity of the packages are identically preserved. The present implementation uses the Message Passing Interface (MPI) library for interprocessor message passing and memory accessing. By choosing an appropriate combination of the available partitioning and coalescing capabilities only during the execution stage, the PENS solver becomes adaptable to different computer architectures from shared-memory to distributed-memory platforms with varying degrees of parallelism. The PENS implementation on the IBM SP2 distributed memory environment at the NASA Ames Research Center obtains 85 percent scalable parallel performance using fine-grain partitioning of single-block CFD domains using up to 128 wide computational nodes. Multi-block CFD simulations of complete aircraft simulations achieve 75 percent perfect load-balanced executions using data coalescing and the two levels of parallelism. SGI PowerChallenge, SGI Origin 2000, and a cluster of workstations are the other platforms where the robustness of the implementation is tested. The performance behavior on the other computer platforms with a variety of realistic problems will be included as this on-going study progresses.

Hatay, Ferhat F.↗

A nonrecursive order N preconditioned conjugate gradient: Range space formulation of MDOF dynamics

While excellent progress has been made in deriving algorithms that are efficient for certain combinations of system topologies and concurrent multiprocessing hardware, several issues must be resolved to incorporate transient simulation in the control design process for large space structures. Specifically, strategies must be developed that are applicable to systems with numerous degrees of freedom. In addition, the algorithms must have a growth potential in that they must also be amenable to implementation on forthcoming parallel system architectures. For mechanical system simulation, this fact implies that algorithms are required that induce parallelism on a fine scale, suitable for the emerging class of highly parallel processors; and transient simulation methods must be automatically load balancing for a wider collection of system topologies and hardware configurations. These problems are addressed by employing a combination range space/preconditioned conjugate gradient formulation of multi-degree-of-freedom dynamics. The method described has several advantages. In a sequential computing environment, the method has the features that: by employing regular ordering of the system connectivity graph, an extremely efficient preconditioner can be derived from the 'range space metric', as opposed to the system coefficient matrix; because of the effectiveness of the preconditioner, preliminary studies indicate that the method can achieve performance rates that depend linearly upon the number of substructures, hence the title 'Order N'; and the method is non-assembling. Furthermore, the approach is promising as a potential parallel processing algorithm in that the method exhibits a fine parallel granularity suitable for a wide collection of combinations of physical system topologies/computer architectures; and the method is easily load balanced among processors, and does not rely upon system topology to induce parallelism.

Kurdila, Andrew J.↗

ESnet/JLab FPGA Accelerated Transport

To increase the science rate for high data rates/volumes, Thomas Jefferson National Accelerator Facility (JLab) has partnered with Energy Sciences Network (ESnet) to define an edge to data center traffic shaping / steering transport capability featuring data event aware network shaping and forwarding. The keystone of this ESnet+JLab FPGA Accelerated Transport (EJFAT) is the joint development of an AI/ML directed dynamic compute work Load Balancer (LB) of UDP streamed data. The LB is a suite consisting of a Field Programmable Gate Array (FPGA) executing the dynamically configurable, low fixed latency LB data plane featuring real-time packet redirection and high throughput, and a control plane running on the FPGA host computer that monitors network and compute farm telemetry in order to make dynamic AI/ML guided decisions for destination compute host redirection/load balancing and destination resource provisioning. The LB provides for three-tier horizontal scaling across LB suites, core compute hosts, and CPUs within a host. The LB effectively provides seamless integration of edge/core computing to support direct experimental data processing for immediate use by JLab science programs and others such as the EIC as well as data centers of the future requiring high throughput and low latency for both hot and cooled data for both running experiment data acquisition systems and data center use cases.

97 MATHEMATICS AND COMPUTING↗

Performance Analysis of Speculative Parallel Adaptive Local Timestepping for Conservation Laws

Stable simulation of conservation laws, such as those used to model fluid dynamics and plasma physics applications, requires the satisfaction of the so-called Courant-Friedrichs-Lewy condition. By allowing regions of the mesh to advance with different timesteps that locally satisfy this stability constraint, significant work reduction can be attained when compared to a time integration scheme using a single timestep size. However, parallelizing this algorithm presents considerable difficulty. Since the stability condition depends on the state of the system, dependencies become dynamic and potentially non-local. In this article, we present an adaptive local timestepping algorithm using an optimistic (Timewarp-based) parallel discrete event simulation. We introduce waiting heuristics to limit misspeculation and a semi-static load balancing scheme to eliminate load imbalance as parts of the mesh require finer or coarser timesteps. Last, we outline an interface for separating the physics of the specific conservation law from the temporal integration allowing for productive adoption of our proposed algorithm. We present a misspeculation study for three conservation laws, demonstrating both the productivity of the local timestepping API, for which 74% of the lines of code are reused across different conservation laws, and the robustness of the waiting heuristics—at most 1.5% of element updates are rolled back. Our performance studies demonstrate up to a 2.8× speedup versus a baseline unoptimized local timestepping approach, a 4x improvement in per-node throughput compared to an MPI parallelization of synchronous timestepping, and scalability up to 3,072 cores on NERSC’s Cori Haswell partition.

97 MATHEMATICS AND COMPUTING↗

Voltage regulator dissipates minimal power and functions as a voltage divider

Regulator requires minimum amount of power for voltage division and it is not required continuously. The only power loss, except for regulating purposes, is that needed to provide for imbalances in load current requirements. For balanced loads, only leakage current flows through regulating transistors.

Hester, H. B.↗

Scaling Demand Flexibility: Building on 30 Years of Energy Efficiency Success

With electricity consumption across the United States (US) and Canada anticipated to grow, energy efficiency program administrators have a key role to play in helping to ensure energy affordability and reliability in support of the broader economic systems utilities and grid support. Connected, demand side load balancing solutions, such as load shifting heating, ventilation and air conditioning (HVAC) systems and managed charging for electric vehicles (EVs), can dynamically manage energy, allowing for more volumetric electricity consumption without incurring the expense of upgraded transmission and distribution capabilities. When combined, or aggregated, many small loads can be managed to have meaningful impact on energy demand on the grid. Utilities and their partners have an opportunity to leverage decades of experience and the infrastructure needed to assess, design, implement, and measure programs to scale up the adoption of equipment with built-in load flexibility capabilities. Current efforts among a wide variety of electricity system service providers, utilities, standards agencies, regulators, national labs and private industry stakeholders aim to identify common standards, metrics, and methodologies for valuing grid services offered by demand side equipment. By combining those efforts with decades of proven energy efficiency resources, utilities are poised to effectuate a scaling up of equipment with energy management capabilities installed in homes and businesses across the US and Canada. This paper will provide an overview of how utilities are approaching this era of load growth and new peak demands across the United States and Canada. It will highlight the specific strategies that program administrators are employing to advance market transformation for grid-enabled products and devices that have the greatest potential to reduce energy use and increase load flexibility.

Grant, Peter↗

Did the GPU obfuscate the load imbalance in my MPI simulation?

The current proliferation of GPU-based HPC systems necessitates a method for assessing the performance of simulations on heterogeneous machines. The addition of GPUs to a system adds multiple hierarchical levels of parallelism to the node architecture. In this paper, we demonstrate that the traditional load imbalance metric is insufficient for capturing the load imbalance on GPU-based machines, since it treats the GPU as a monolithic entity and ignores the internal parallelism. We propose a new hierarchical metric that improves the correlation of measured performance and application workload by up to 20.61%. Using our metric for determining application load instead of the traditional metric as the input for the load balancing algorithm reduces the residual load imbalance by up to 4× in our application.

Eberius, David↗

Prediction Interval Development for Wind-Tunnel Balance Check-Loading

Results from the Facility Analysis Verification and Operational Reliability project revealed a critical gap in capability in ground-based aeronautics research applications. Without a standardized process for check-loading the wind-tunnel balance or the model system, the quality of the aerodynamic force data collected varied significantly between facilities. A prediction interval is required in order to confirm a check-loading. The prediction interval provides an expected upper and lower bound on balance load prediction at a given confidence level. A method has been developed which accounts for sources of variability due to calibration and check-load application. The prediction interval method of calculation and a case study demonstrating its use is provided. Validation of the methods is demonstrated for the case study based on the probability of capture of confirmation points.

Landman, Drew↗

A Universal Algorithm for the Detection of Bi-directional Gage Output Characteristics

A universal algorithm was developed that may be used to assess the bi-directional characteristics of the gage outputs of a wind tunnel strain-gage balance. The algorithm assumes that balance loads and gage outputs are described in the design format of the balance. It can also be applied to balance calibration data that is processed by using either the Iterative Method or the Non-Iterative Method. The algorithm uses an estimate of the bi-directional part of a gage output at load capacity as input. In addition, the statistical significance of the principle absolute value term in the regression model of either the gage output or the related primary load component is determined. A gage output is assumed to be bi-directional if two conditions are fulfilled: the bi-directional part of the output at load capacity exceeds 0.5 percent of the maximum output at load capacity; the p-value of the principle absolute value term of the regression model of the balance data is less than the threshold of 0.001. Data from the calibration of two six-component force balances and one five-component semi-span balance are used to illustrate the application of the universal detection algorithm.

wind tunnel test↗