Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “vectorizability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Vectorizable multigrid algorithms for transonic flow calculations

The analysis and incorporation into a multigrid scheme of several vectorizable algorithms are discussed. Von Neumann analyses of vertical line, horizontal line, and alternating direction ZEBRA algorithms were performed; and the results were used to predict their multigrid damping rates. The algorithms were then successfully implemented in a transonic conservative full-potential computer program. The convergence acceleration effect of multiple grids is shown and the convergence rates of the vectorizable algorithms are compared to the convergence rates of standard successive line overrelaxation (SLOR) algorithms.

Melson, N. D.↗

Vectorizable algorithms for adaptive schemes for rapid analysis of SSME flows

An initial study into vectorizable algorithms for use in adaptive schemes for various types of boundary value problems is described. The focus is on two key aspects of adaptive computational methods which are crucial in the use of such methods (for complex flow simulations such as those in the Space Shuttle Main Engine): the adaptive scheme itself and the applicability of element-by-element matrix computations in a vectorizable format for rapid calculations in adaptive mesh procedures.

Oden, J. Tinsley↗

Vectorizable multigrid algorithms for transonic-flow calculations

The analysis and the incorporation into a multigrid scheme of several vectorizable algorithms are discussed. von Neumann analyses of vertical-line, horizontal-line, and alternating-direction ZEBRA algorithms were performed; and the results were used to predict their multigrid damping rates. The algorithms were then successfully implemented in a transonic conservative full-potential computer program. The convergence acceleration effect of multiple grids is shown, and the convergence rates of the vectorizable algorithms are compared with those of standard successive-line overrelaxation (SLOR) algorithms.

Melson, N. D.↗

VecPAC: A Vectorizable and Precision-Aware CGRA

This paper proposes VecPAC -- a vectorizable and precision-aware coarse-grained reconfigurable array (CGRA) design. VecPAC integrates CGRA tiles with scalar functional units and specialized tiles with vector functional units that can trade off the number of vector lanes for the accuracy of the computation. We discuss the architecture design and present the related compilation framework. The experimental evaluation on a set of applications from three different domains (embedded, machine learning, and high-performance computing) shows that the hybrid design of VecPAC outperforms CGRAs with only scalar functional units by 1.48x, while providing higher scalability.

Tan, Cheng↗

Guidelines for developing vectorizable computer programs

Some fundamental principles for developing computer programs which are compatible with array-oriented computers are presented. The emphasis is on basic techniques for structuring computer codes which are applicable in FORTRAN and do not require a special programming language or exact a significant penalty on a scalar computer. Researchers who are using numerical techniques to solve problems in engineering can apply these basic principles and thus develop transportable computer programs (in FORTRAN) which contain much vectorizable code. The vector architecture of the ASC is discussed so that the requirements of array processing can be better appreciated. The "vectorization" of a finite-difference viscous shock-layer code is used as an example to illustrate the benefits and some of the difficulties involved. Increases in computing speed with vectorization are illustrated with results from the viscous shock-layer code and from a finite-element shock tube code. The applicability of these principles was substantiated through running programs on other computers with array-associated computing characteristics, such as the Hewlett-Packard (H-P) 1000-F.

Miner, E. W.↗

Vectorizable implicit algorithms for the flux-difference split, three-dimensional Navier-Stokes equations

The computational efficiency of four vectorizable implicit algorithms is assessed when applied to calculate steady-state solutions to the three-dimensional, incompressible Navier-Stokes equations in general coordinates. Two of these algorithms are characterized as hybrid schemes; that is, they combine some approximate factorization in two coordinate directions with relaxation in the remaining spatial direction. The other two algorithms utilize an approximate factorization approach which yields two-factor algorithms for three-dimensional systems. All four algorithms are implemented in identical high-resolution upwind schemes for the flux-difference split Navier-Stokes equations. These highly nonlinear schemes are obtained by extending an implicit Total Variation Diminishing (TVD) scheme recently developed for linear one-dimensional systems of hyperbolic conservation laws to the three-dimensional Navier-Stokes equations. The computation of vortical flow over a sharp-edged, thin delta wing has been chosen as a common numerical test case. The convergence of the algorithms is discussed and the accuracy of the computed flow-field results is assessed. The validity of the present results are demonstrated by a comparison with experimental data.

Hartwich, P. M.↗

A novel MC TRT method: vectorizable variance reduction for the energy spectra

Data produced up to now gives us confidence that we will continue to see variance reduction behavior at amenable runtimes as we move this test bed forward. This method has characteristics of event and history based Monte Carlo transport, without the bulky modifications required by event based Monte Carlo while avoiding some overhead.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Implicit, vectorizable schemes for the flux-difference split, three-dimensional Navier-Stokes equations

Two hybrid upwind models are defined for solving the Euler equations. The algorithms both employ approximate factorization (AF) in crossplane and symmetric block Gauss-Seidel relaxation in the third direction. One approach adds an additional factorization step to lower the number of required grid point operations for inversion of the block tridiagonal matrices; however, the move permits only one third of the operations to be vectorized. Finite difference solutions are calculated on a C-H-type grid, in this case enveloping a slender, sharp-edged delta wing. Sample data are provided for the calculated vortex flow for Re of 10,000, at a 20.5 deg angle of attack, represented in a crossflow velocity vector plot and in a spanwise pressure coefficient distribution. The AF scheme, without additional factorization, when used with a grid covering 51 x 51 x 72 points provides a convergent solution with no time step lasting longer than 0.00001 sec.

Liu, C. H.↗

On the suitability of the connection machine for direct particle simulation

The algorithmic structure was examined of the vectorizable Stanford particle simulation (SPS) method and the structure is reformulated in data parallel form. Some of the SPS algorithms can be directly translated to data parallel, but several of the vectorizable algorithms have no direct data parallel equivalent. This requires the development of new, strictly data parallel algorithms. In particular, a new sorting algorithm is developed to identify collision candidates in the simulation and a master/slave algorithm is developed to minimize communication cost in large table look up. Validation of the method is undertaken through test calculations for thermal relaxation of a gas, shock wave profiles, and shock reflection from a stationary wall. A qualitative measure is provided of the performance of the Connection Machine for direct particle simulation. The massively parallel architecture of the Connection Machine is found quite suitable for this type of calculation. However, there are difficulties in taking full advantage of this architecture because of lack of a broad based tradition of data parallel programming. An important outcome of this work has been new data parallel algorithms specifically of use for direct particle simulation but which also expand the data parallel diction.

Dagum, Leonard↗

Application of the multigrid method to grid generation

The multigrid method (MGM), used to numerically solve the pair of nonlinear elliptic equations commonly used to generate two dimensional boundary-fitted coordinate systems is discussed. Two different geometries are considered: one involving a coordinate system fitted about a circle and the other selected for an impinging jet flow problem. Two different relaxation schemes are tried: one is successive point overrelaxation and the other is a four-color scheme vectorizeable to take advantage of a parallel processor computer for greater computational speed. Results using MGM are compared with those using SOR (doing successive overrelaxations with the corresponding relaxation scheme on the fine grid only). It is found that MGM becomes significantly more effective than SOR as more accuracy is demanded and as more corrective grids, or more grid points, are used. For the accuracy required, it is found that MGM is two to three times faster than SOR in computing time. With the four-color relaxation scheme as applied to the impinging jet problem, the advantage of MGM over SOR is not as great. This may be due to the effect of a poor initial guess on MGM for this problem.

Ohring, S.↗

Experiences in using the CYBER 203 for three-dimensional transonic flow calculations

In this paper, the authors report on some of their experiences modifying two three-dimensional transonic flow programs (FLO22 and FLO27) for use on the NASA Langley Research Center CYBER 203. Both of the programs discussed were originally written for use on serial machines. Several methods were attempted to optimize the execution of the two programs on the vector machine, including: (1) leaving the program in a scalar form (i.e., serial computation) with compiler software used to optimize and vectorize the program, (2) vectorizing parts of the existing algorithm in the program, and (3) incorporating a new vectorizable algorithm (ZEBRA I or ZEBRA II) in the program.

Melson, N. D.↗

Use of CYBER 203 and CYBER 205 computers for three-dimensional transonic flow calculations

Experiences are discussed for modifying two three-dimensional transonic flow computer programs (FLO 22 and FLO 27) for use on the CDC CYBER 203 computer system. Both programs were originally written for use on serial machines. Several methods were attempted to optimize the execution of the two programs on the vector machine: leaving the program in a scalar form (i.e., serial computation) with compiler software used to optimize and vectorize the program, vectorizing parts of the existing algorithm in the program, and incorporating a vectorizable algorithm (ZEBRA I or ZEBRA II) in the program. Comparison runs of the programs were made on CDC CYBER 175. CYBER 203, and two pipe CDC CYBER 205 computer systems.

Melson, N. D.↗

Efficient solution of the Euler and Navier-Stokes equations with a vectorized multiple-grid algorithm

A multiple-grid algorithm for use in efficiently obtaining steady solutions to the Euler and Navier-Stokes equations is presented. The convergence of the explicit MacCormack algorithm on a fine grid is accelerated by propagating transients from the domain using a sequence of successively coarser grids. Both the fine and coarse grid schemes are readily vectorizable. The combination of multiple-gridding and vectorization results in substantially reduced computational times for the numerical solution of a wide range of flow problems. Results are presented for subsonic, transonic, and supersonic inviscid flows and for subsonic attached and separated laminar viscous flows. Work reduction factors over a scalar, single-grid algorithm range as high as 76.8.

Navier Stokes simulation↗

Improved relaxation schemes for transonic potential calculations

A block relaxation scheme, grouped in a red-black ordering, is applied to transonic airfoil calculations using body fitted coordinates. The scheme is simple and is easily vectorizable. Detailed comparisons with Approximate Factorization Method (AF2) are presented and it is shown that the improved relaxation scheme is competitive in all cases considered. Transonic results, of engineering accuracy, on an 0-type grid of 149 x 30 points, are ususally obtained within two hundred iterations (approximately 40 seconds on Cyber 175).

Hafez, M.↗

Efficient solution of the Euler and Navier-Stokes equations with a vectorized multiple-grid algorithm

A multiple-grid algorithm for use in efficiently obtaining steady solutions to the Euler and Navier-Stokes equations is presented. The convergence of the explicit MacCormack algorithm on a fine grid is accelerated by propagating transients from the domain using a sequence of successively coarser grids. Both the fine and coarse grid schemes are readily vectorizable. The combination of multiple-gridding and vectorization results in substantially reduced computational times for the numerical solution of a wide range of flow problems. Results are presented for subsonic, transonic, and supersonic inviscid flows and for subsonic attached and separated laminar viscous flows. Work reduction factors over a scalar, single-grid algorithm range as high as 76.8. Previously announced in STAR as N83-24467

Chima, R. V.↗

A fully vectorized numerical solution of the incompressible Navier-Stokes equations

A vectorizable algorithm is presented for the implicit finite difference solution of the incompressible Navier-Stokes equations in general curvilinear coordinates. The unsteady Reynolds averaged Navier-Stokes equations solved are in two dimension and non-conservative primitive variable form. A two-layer algebraic eddy viscosity turbulence model is used to incorporate the effects of turbulence. Two momentum equations and a Poisson pressure equation, which is obtained by taking the divergence of the momentum equations and satisfying the continuity equation, are solved simultaneously at each time step. An elliptic grid generation approach is used to generate a boundary conforming coordinate system about an airfoil. The governing equations are expressed in terms of the curvilinear coordinates and are solved on a uniform rectangular computational domain. A checkerboard SOR, which can effectively utilize the computer architectural concept of vector processing, is used for iterative solution of the governing equations.

Patel, N.↗

A new algorithm for tuning of computed radiances for HIRS2/MSU

Small biases of the order of 1 C exist in brightness temperatures computed for a number of atmospheric sounding channels using radiosonde reports of atmospheric temperature humidity profile compared to those of collocated HIRS2/MSU observations on TIROS N. These biases are attributed to errors in the computed atmospheric transmittances functions. Channel dependent empirical tuning coefficients were found such that the biases in the channel brightness temperatures are removed if the transmittances used to calculate these brightness temperatures are modified. Possible shortcomings of this method are that some of the bias errors may be due to instrumental calibration problems and that the part that is computational may not be of the form assumed in the equation used. Form of tuning was implemented in the calculation which has the potential of distinguishing between calibration and calculation errors and is also computationally faster and more easily vectorizable.

Susskind, J.↗