Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “load balancing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

Cluster Computation of Flight Reynolds Number Flows

The performance of a workstation cluster used for the solution of the Reynolds-averaged Navier-Stokes equations is compared with a conventional vector supercomputer architecture. The application simulation of the steady flowfield about a transonic transport was computed using an implicit diagonal scheme in an overset mesh framework. Static load balancing was used, while coarse grain decomposition was achieved by solution of a grid zone per processor. Price/performance ratios are estimated for several scenarios in which such clusters may be utilized.

Atwood, Christopher A.↗

Parallel and Distributed Computational Fluid Dynamics: Experimental Results and Challenges

This paper describes several results of parallel and distributed computing using a large scale production flow solver program. A coarse grained parallelization based on clustering of discretization grids combined with partitioning of large grids for load balancing is presented. An assessment is given of its performance on distributed and distributed-shared memory platforms using large scale scientific problems. An experiment with this solver, adapted to a Wide Area Network execution environment is presented. We also give a comparative performance assessment of computation and communication times on both the tightly and loosely-coupled machines.

Djomehri, Mohammad Jahed↗

Parallel Programming Strategies for Irregular Adaptive Applications

Achieving scalable performance for dynamic irregular applications is eminently challenging. Traditional message-passing approaches have been making steady progress towards this goal; however, they suffer from complex implementation requirements. The use of a global address space greatly simplifies the programming task, but can degrade the performance for such computations. In this work, we examine two typical irregular adaptive applications, Dynamic Remeshing and N-Body, under competing programming methodologies and across various parallel architectures. The Dynamic Remeshing application simulates flow over an airfoil, and refines localized regions of the underlying unstructured mesh. The N-Body experiment models two neighboring Plummer galaxies that are about to undergo a merger. Both problems demonstrate dramatic changes in processor workloads and interprocessor communication with time; thus, dynamic load balancing is a required component.

Biswas, Rupak↗

Platform-Independence and Scheduling In a Multi-Threaded Real-Time Simulation

Aviation research often relies on real-time, pilot-in-the-loop flight simulation as a means to develop new flight software, flight hardware, or pilot procedures. Often these simulations become so complex that a single processor is incapable of performing the necessary computations within a fixed time-step. Threads are an elegant means to distribute the computational work-load when running on a symmetric multi-processor machine. However, programming with threads often requires operating system specific calls that reduce code portability and maintainability. While a multi-threaded simulation allows a significant increase in the simulation complexity, it also increases the workload of a simulation operator by requiring that the operator determine which models run on which thread. To address these concerns an object-oriented design was implemented in the NASA Langley Standard Real-Time Simulation in C++ (LaSRS++) application framework. The design provides a portable and maintainable means to use threads and also provides a mechanism to automatically load balance the simulation models.

Sugden, Paul P.↗

Navier-Stokes Aerodynamic Simulation of the V-22 Osprey on the Intel Paragon MPP

The paper will describe the Development of a general three-dimensional multiple grid zone Navier-Stokes flowfield simulation program (ENS3D-MPP) designed for efficient execution on the Intel Paragon Massively Parallel Processor (MPP) supercomputer, and the subsequent application of this method to the prediction of the viscous flowfield about the V-22 Osprey tiltrotor vehicle. The flowfield simulation code solves the thin Layer or full Navier-Stoke's equation - for viscous flow modeling, or the Euler equations for inviscid flow modeling on a structured multi-zone mesh. In the present paper only viscous simulations will be shown. The governing difference equations are solved using a time marching implicit approximate factorization method with either TVD upwind or central differencing used for the convective terms and central differencing used for the viscous diffusion terms. Steady state or Lime accurate solutions can be calculated. The present paper will focus on steady state applications, although time accurate solution analysis is the ultimate goal of this effort. Laminar viscosity is calculated using Sutherland's law and the Baldwin-Lomax two layer algebraic turbulence model is used to compute the eddy viscosity. The Simulation method uses an arbitrary block, curvilinear grid topology. An automatic grid adaption scheme is incorporated which concentrates grid points in high density gradient regions. A variety of user-specified boundary conditions are available. This paper will present the application of the scalable and superscalable versions to the steady state viscous flow analysis of the V-22 Osprey using a multiple zone global mesh. The mesh consists of a series of sheared cartesian grid blocks with polar grids embedded within to better simulate the wing tip mounted nacelle. MPP solutions will be shown in comparison to equivalent Cray C-90 results and also in comparison to experimental data. Discussions on meshing considerations, wall clock execution time, load balancing, and scalability will be provided.

Vadyak, Joseph↗

A Unique Power System For The ISS Fluids And Combustion Facility

Unique power control technology has been incorporated into an electrical power control unit (EPCU) for the Fluids and Combustion Facility (FCF). The objective is to maximize science throughput by providing a flexible power system that is easily reconfigured by the science payload. Electrical power is at a premium on the International Space Station (ISS). The EPCU utilizes advanced power management techniques to maximize the power available to the FCF experiments. The EPCU architecture enables dynamic allocation of power from two ISS power channels for experiments. Because of the unique flexible remote power controller (FRPC) design, power channels can be paralleled while maintaining balanced load sharing between the channels. With an integrated and redundant architecture, the EPCU can tolerate multiple faults and still maintain FCF operation. It is important to take full advantage of the EPCU functionality. The EPCU acts as a buffer between the experimenter and the ISS power system with all its complex requirements. However, FCF science payload developers will still need to follow guidelines when designing the FCF payload power system. This is necessary to ensure power system stability, fault coordination, electromagnetic compatibility, and maximum use of available power for gathering scientific data.

Fox, David A.↗

RANS-MP: A Portable Parallel Navier-Stokes Solver

RANS-MP, a new implementation of a single-grid Navier-Stokes solver using the diagonalized Beam-Warming approximate-factorization scheme, is presented. This first release of the completely rewritten solver employs the following optimizations: (1) Bi-directional multi-partition method for the ADI solver part; this improves granularity and load balance; (2) Improved cache usage through elimination of non-unit-stride array access (possible in part due to multi-partitioning); (3) Preprocessing of communicating boundary conditions to streamline logic during time stepping; (4) Truly parallel, high-performance I/O using the newly-developed MPI-IO library; (5) Elimination of large amounts of redundant operations through efficient use of workspace. Results of some realistic wing computations on the IBM SP2 computer will be presented. We will demonstrate that excellent absolute performance and scalability are obtained with RANS-MP, even for relatively small grid sizes. Besides high performance, an outstanding feature of RANS-MP is its true portability, due to the use of the portable message passing and I/O libraries MPI and MPI-IO.

VanderWijngaart, Rob F.↗

A Portable MPI Implementation of the SPAI Preconditioner in ISIS++

A parallel MPI implementation of the Sparse Approximate Inverse (SPAI) preconditioner is described. SPAI has proven to be a highly effective preconditioner, and is inherently parallel because it computes columns (or rows) of the preconditioning matrix independently. However, there are several problems that must be addressed for an efficient MPI implementation: load balance, latency hiding, and the need for one-sided communication. The effectiveness, efficiency, and scaling behavior of our implementation will be shown for different platforms.

Barnard, Stephen T.↗

A Fast-Time Study of Aircraft Reordering in Arrival Sequencing and Scheduling

In order to ensure that the safe capacity of the terminal area is not exceeded, Air Traffic Management ATM often places restrictions on arriving flights transitioning from en route airspace to terminal airspace. This restriction of arrival traffic is commonly referred to as arrival flow management, and includes techniques such as metering, vectoring, fix-load balancing, and the imposition of miles-in-trail separations. These restrictions are enacted without regard for the relative priority which airlines may be placing on individual flights based on factors such as crew criticality, passenger connectivity, critical turn times, gate availability, on-time performance, fuel status, or runway preference. The development of new arrival flow management techniques which take into consideration priorities expressed by air carriers will likely reduce the economic impact of ATM restrictions on the airlines and lead to increased airline economic efficiency by allowing airlines to have greater control over their individual arrival banks of aircraft. NASA and the Federal Aviation Administration (FAA) have designed and developed a suite of software decision support tools (DSTs) collectively known as the Center TRACON Automation System (CTAS). One of these tools, the Traffic Management Advisor (TMA) is currently being used at the Fort Worth Air Route Traffic Control Center to perform arrival flow management of traffic into the Dallas/Fort Worth airport (DFW). The TMA is a time-based strategic planning tool that assists Traffic Management Coordinators (TMCs) and En Route Air Traffic Controllers in efficiently balancing arrival demand with airport capacity. The primary algorithm in the TMA is a real-time scheduler which generates efficient landing sequences and landing times for arrivals within about 200 no a. from touchdown. This scheduler will sequence aircraft so that they arrive in a first- come - first-served (FCFS) order. While FCFS sequencing establishes a fair order based on estimated times of arrival, it does not take into account individual airline priorities among incoming flights. NASA is exploring the possibility of allowing airlines to express relative arrival priorities to air traffic management through the development of new CTAS scheduling algorithms which take into consideration airline arrival preferences. The accommodation of airline priorities in arrival sequencing and scheduling would under most circumstances result in a deviation from a "natural" or FCFS arrival order. As a First step toward developing airline influenced sequencing algorithms, an investigation was conducted to determine the feasibility of reordering arrival traffic from a strict FCFS sequence. A fast-time simulation has been developed which allows statistical evaluation of sequencing and scheduling algorithms for arrival traffic at the Dallas/Fort Worth Airport. In contrast to real-time simulation or field tests, which would require on the order of ninety minutes to examine a single traffic rush period, the fast-time simulation allows examination of multiple rush periods in a matter of seconds.

Carr, Greg↗

Communication Improvement for the LU NAS Parallel Benchmark: A Model for Efficient Parallel Relaxation Schemes

The first release of the MPI version of the LU NAS Parallel Benchmark (NPB2.0) performed poorly compared to its companion NPB2.0 codes. The later LU release (NPB2.1 & 2.2) runs up to two and a half times faster, thanks to a revised point access scheme and related communications scheme. The new scheme sends substantially fewer messages. is cache "friendly", and has a better load balance. We detail the, observations and modifications that resulted in this efficiency improvement, and show that the poor behavior of the original code resulted from deriving a message passing scheme from an algorithm originally devised for a vector architecture.

Yarrow, Maurice↗

Self-Avoiding Walks over Adaptive Triangular Grids

In this paper, we present a new approach to constructing a "self-avoiding" walk through a triangular mesh. Unlike the popular approach of visiting mesh elements using space-filling curves which is based on a geometric embedding, our approach is combinatorial in the sense that it uses the mesh connectivity only. We present an algorithm for constructing a self-avoiding walk which can be applied to any unstructured triangular mesh. The complexity of the algorithm is O(n x log(n)), where n is the number of triangles in the mesh. We show that for hierarchical adaptive meshes, the algorithm can be easily parallelized by taking advantage of the regularity of the refinement rules. The proposed approach should be very useful in the run-time partitioning and load balancing of adaptive unstructured grids.

Heber, Gerd↗

Recent Progress on the Parallel Implementation of Moving-Body Overset Grid Schemes

Viscous calculations about geometrically complex bodies in which there is relative motion between component parts is one of the most computationally demanding problems facing CFD researchers today. This presentation documents results from the first two years of a CHSSI-funded effort within the U.S. Army AFDD to develop scalable dynamic overset grid methods for unsteady viscous calculations with moving-body problems. The first pan of the presentation will focus on results from OVERFLOW-D1, a parallelized moving-body overset grid scheme that employs traditional Chimera methodology. The two processes that dominate the cost of such problems are the flow solution on each component and the intergrid connectivity solution. Parallel implementations of the OVERFLOW flow solver and DCF3D connectivity software are coupled with a proposed two-part static-dynamic load balancing scheme and tested on the IBM SP and Cray T3E multi-processors. The second part of the presentation will cover some recent results from OVERFLOW-D2, a new flow solver that employs Cartesian grids with various levels of refinement, facilitating solution adaption. A study of the parallel performance of the scheme on large distributed- memory multiprocessor computer architectures will be reported.

Wissink, Andrew↗

Self-Avoiding Walks Over Adaptive Triangular Grids

Space-filling curves is a popular approach based on a geometric embedding for linearizing computational meshes. We present a new O(n log n) combinatorial algorithm for constructing a self avoiding walk through a two dimensional mesh containing n triangles. We show that for hierarchical adaptive meshes, the algorithm can be locally adapted and easily parallelized by taking advantage of the regularity of the refinement rules. The proposed approach should be very useful in the runtime partitioning and load balancing of adaptive unstructured grids.

Heber, Gerd↗

Parallel Computing Strategies for Irregular Algorithms

Parallel computing promises several orders of magnitude increase in our ability to solve realistic computationally-intensive problems, but relies on their efficient mapping and execution on large-scale multiprocessor architectures. Unfortunately, many important applications are irregular and dynamic in nature, making their effective parallel implementation a daunting task. Moreover, with the proliferation of parallel architectures and programming paradigms, the typical scientist is faced with a plethora of questions that must be answered in order to obtain an acceptable parallel implementation of the solution algorithm. In this paper, we consider three representative irregular applications: unstructured remeshing, sparse matrix computations, and N-body problems, and parallelize them using various popular programming paradigms on a wide spectrum of computer platforms ranging from state-of-the-art supercomputers to PC clusters. We present the underlying problems, the solution algorithms, and the parallel implementation strategies. Smart load-balancing, partitioning, and ordering techniques are used to enhance parallel performance. Overall results demonstrate the complexity of efficiently parallelizing irregular algorithms.

Biswas, Rupak↗

An Analysis of Performance Enhancement Techniques for Overset Grid Applications

The overset grid methodology has significantly reduced time-to-solution of high-fidelity computational fluid dynamics (CFD) simulations about complex aerospace configurations. The solution process resolves the geometrical complexity of the problem domain by using separately generated but overlapping structured discretization grids that periodically exchange information through interpolation. However, high performance computations of such large-scale realistic applications must be handled efficiently on state-of-the-art parallel supercomputers. This paper analyzes the effects of various performance enhancement techniques on the parallel efficiency of an overset grid Navier-Stokes CFD application running on an SGI Origin2000 machine. Specifically, the role of asynchronous communication, grid splitting, and grid grouping strategies are presented and discussed. Results indicate that performance depends critically on the level of latency hiding and the quality of load balancing across the processors.

Djomehri, J. J.↗

Measurement of Unsteady Pressure Data on a Large HSCT Semispan Wing and Comparison with Analysis

Experimental data from wind-tunnel tests of the Rigid Semispan Model (RSM) performed at NASA Langley's Transonic Dynamics Tunnel (TDT) are presented. The primary focus of the paper is on data obtained from testing of the RSM on the Oscillating Turntable (OTT). The OTT is capable of oscillating models in pitch at various amplitudes and frequencies about mean angles of attack. Steady and unsteady pressure data obtained during testing of the RSM on the OTT is presented and compared to data obtained from previous tests of the RSM on a load balance and on a Pitch and Plunge Apparatus (PAPA). Testing of the RSM on the PAPA resulted in utter boundaries that were strongly dependent on angle of attack across the Mach number range. Pressure data from all three tests indicates the existence of vortical flows at moderate angles of attack. The correlation between the vortical flows and the unusual utter boundaries from the RSM/PAPA test is discussed. Comparisons of experimental data with analyses using the CFL3Dv6 computational fluid dynamics code are presented.

Scott, Robert C.↗

Scalability of a Low-Cost Multi-Teraflop Linux Cluster for High-End Classical Atomistic and Quantum Mechanical Simulations

Scalability of a low-cost, Intel Xeon-based, multi-Teraflop Linux cluster is tested for two high-end scientific applications: Classical atomistic simulation based on the molecular dynamics method and quantum mechanical calculation based on the density functional theory. These scalable parallel applications use space-time multiresolution algorithms and feature computational-space decomposition, wavelet-based adaptive load balancing, and spacefilling-curve-based data compression for scalable I/O. Comparative performance tests are performed on a 1,024-processor Linux cluster and a conventional higher-end parallel supercomputer, 1,184-processor IBM SP4. The results show that the performance of the Linux cluster is comparable to that of the SP4. We also study various effects, such as the sharing of memory and L2 cache among processors, on the performance.

Kikuchi, Hideaki↗

Chapman-1024 Processor Shared Memory

NASA has developed new technology that improves upon weakness in current mainstream supercomputer designs: those of "scalability," "humadmachine interface," and "load balancing." The system simplifies running large computer simulations of national and international importance like climate prediction and space vehicle design.

Ciotii, Robert↗