Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel coordinates”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Parallel CFD Algorithms for Aerodynamical Flow Solvers on Unstructured Meshes

The Advisory Group for Aerospace Research and Development (AGARD) has requested my participation in the lecture series entitled Parallel Computing in Computational Fluid Dynamics to be held at the von Karman Institute in Brussels, Belgium on May 15-19, 1995. In addition, a request has been made from the US Coordinator for AGARD at the Pentagon for NASA Ames to hold a repetition of the lecture series on October 16-20, 1995. I have been asked to be a local coordinator for the Ames event. All AGARD lecture series events have attendance limited to NATO allied countries. A brief of the lecture series is provided in the attached enclosure. Specifically, I have been asked to give two lectures of approximately 75 minutes each on the subject of parallel solution techniques for the fluid flow equations on unstructured meshes. The title of my lectures is "Parallel CFD Algorithms for Aerodynamical Flow Solvers on Unstructured Meshes" (Parts I-II). The contents of these lectures will be largely review in nature and will draw upon previously published work in this area. Topics of my lectures will include: (1) Mesh partitioning algorithms. Recursive techniques based on coordinate bisection, Cuthill-McKee level structures, and spectral bisection. (2) Newton's method for large scale CFD problems. Size and complexity estimates for Newton's method, modifications for insuring global convergence. (3) Techniques for constructing the Jacobian matrix. Analytic and numerical techniques for Jacobian matrix-vector products, constructing the transposed matrix, extensions to optimization and homotopy theories. (4) Iterative solution algorithms. Practical experience with GIVIRES and BICG-STAB matrix solvers. (5) Parallel matrix preconditioning. Incomplete Lower-Upper (ILU) factorization, domain-decomposed ILU, approximate Schur complement strategies.

Barth, Timothy J.↗

Wave-interactions in supersonic and hypersonic flows

Work completed under the current grant comprises the start of a theoretical and computational attack on the subharmonic route to secondary instabilities in compressible flows. The total flow field in this problem is made up of the following components: (1) a steady streamwise mean boundary layer flow which depends only on the normal space component y; (2) a two-dimensional time dependent T-S wave which moves with wavespeed c and has no spanwise dependence; and (3) a fully three-dimensional, time dependent T-S wave whose streamwise wavenumber is half of the streamwise wavenumber associated with the two-dimensional T-S wave in b. If a frame of reference is adopted which moves with the wavespeed c of the 2-D T-S wave, the time dependence of this portion of the flow can be eliminated. The effective steady mean flow in this problem is now the sum of the original parallel steady mean flow and the initial 2-D T-S instability. Dependence on the streamwise coordinate x in this mean flow can be extracted by assuming normal mode expansions involving complex exponentials and the streamwise wavenumber a. However, it is important to note that, because this is a wave-wave interaction problem, unlike the usual linear instability case, both the complex exponential, and its complex conjugate, must be retained in describing the 2-D T-S wave. The role of the perturbation to the steady mean flow is now played by the 3-D time dependent T-S wave. In treating this wave, normal modes in the streamwise and spanwise directions and time may be used. Consistent with the subharmonic nature of this transition route, the streamwise wavenumber is a/2, and complex conjugates of the complex exponential must be employed. This is not the case with the modes giving z and t dependence with wavespeed o and spanwise wavenumber B as the effective mean flow quantities are independent of z and their time dependence is accounted for by the moving frame of reference. Consequently, the wave-wave interaction which will produce mean flow modification occurs only through the streamwise exponentials.

Lakin, William D.↗

Framework for Extensible, Asynchronous Task Scheduling (FEATS) in Fortran

Most parallel scientific programs contain compiler directives (pragmas) such as those from OpenMP, explicit calls to runtime library procedures such as those implementing the Message Passing Interface (MPI), or compiler-specific language extensions such as those provided by CUDA. By contrast, the recent Fortran standards empower developers to express parallel algorithms without directly referencing lower-level parallel programming models. Fortran’s parallel features place the language within the Partitioned Global Address Space (PGAS) class of programming models. When writing programs that exploit data-parallelism, application developers often find it straightforward to develop custom parallel algorithms. Problems involving complex, heterogeneous, staged calculations, however, pose much greater challenges. Such applications require careful coordination of tasks in a manner that respects dependencies prescribed by a directed acyclic graph. When rolling one’s own solution proves difficult, extending a customizable framework becomes attractive. The paper presents the design, implementation, and use of the Framework for Extensible Asynchronous Task Scheduling (FEATS), which we believe to be the first task-scheduling tool written in modern Fortran. We describe the benefits and compromises associated with choosing Fortran as the implementation language, and we propose ways in which future Fortran standards can best support the use case in this paper.

Modern Fortran↗

An MPA-IO interface to HPSS

This paper describes an implementation of the proposed MPI-IO (Message Passing Interface - Input/Output) standard for parallel I/O. Our system uses third-party transfer to move data over an external network between the processors where it is used and the I/O devices where it resides. Data travels directly from source to destination, without the need for shuffling it among processors or funneling it through a central node. Our distributed server model lets multiple compute nodes share the burden of coordinating data transfers. The system is built on the High Performance Storage System (HPSS), and a prototype version runs on a Meiko CS-2 parallel computer.

Jones, Terry↗

An MPI-IO Interface to HPSS

This paper describes an implementation of the proposed MPI-IO standard for parallel I/O. Our system uses third-party transfer to move data over an external network between the processors where it is used and the I/O devices where it resides. Data travels directly from source to destination, without the need for shuffling it among processors or funneling it through a central node. Our distributed server model lets multiple compute nodes share the burden of coordinating data transfers. The system is built on the High Performance Storage System (HPSS), and a prototype version runs on a Meiko CS-2 parallel computer.

Parallel Processing↗

Quantization and symmetry in periodic coverage patterns with applications to earth observation

Analytical approaches based on an idealized physical model and concepts from number theory show that in periodic coverage patterns, uniquely defined by their revolution numbers R (orbital) and N (rotational), the subnodal points are earth-fixed, and they divide the equator into R equal segments of length s. The ascending subsatellite trace crosses each point once (only) each period. The descending subnodal points coincide with the ascending points if the integers N and R have like parity, and bisect the intervals between them if opposite. The interval between consecutive unidirectional crossings is Ns. Symmetries extend the equatorial results to all parallels of latitude. Complete periodic patterns of traces exhibit an overall symmetry, with trace intersections confined to discrete coordinate values which are quantized in longitude (basic s-unit) and symmetric in latitude.

King, J. C.↗

Geometric registration and rectification of spaceborne SAR imagery

This paper describes the development of automated location and geometric rectification techniques for digitally processed synthetic aperture radar (SAR) imagery. A software package has been developed that is capable of determining the absolute location of an image pixel to within 60 m using only the spacecraft ephemeris data and the characteristics of the SAR data collection and processing system. Based on this location capability algorithms have been developed that geometrically rectify the imagery, register it to a common coordinate system and mosaic multiple frames to form extended digital SAR maps. These algorithms have been optimized using parallel processing techniques to minimize the operating time. Test results are given using Seasat SAR data.

Curlander, J. C.↗

Hybrid solid element with a traction-free cylindrical surface

An eight node solid element with two parallel faces and one traction-free cylindrical surface is derived using the assumed stress hybrid method. Cylindrical coordinates are used so that the assumed stresses satisfy the equilibrium equations as well as the traction-free condition over the cylindrical boundary. In the limiting case of plane stress conditions the assumed stresses also satisfy the compatibility conditions. Example solutions have demonstrated the advantage of using this special element for analyzing solids with circular holes.

Pian, T. H. H.↗

Recognizing Patterns In Log-Polar Coordinates

Log-Hough transform is basis of improved method for recognition of patterns - particularly, straight lines - in noisy images. Takes advantage of rotational and scale invariance of mapping from Cartesian to log-polar coordinates, and offers economy of representation and computation. Unification of iconic and Hough domains simplifies computations in recognition and eliminates erroneous quantization of slopes attributable to finite spacing of Cartesian coordinate grid of classical Hough transform. Equally efficient recognizing curves. Log-Hough transform more amenable to massively parallel computing architectures than traditional Cartesian Hough transform. "In-place" nature makes it possible to apply local pixel-neighborhood processing.

Weiman, Carl F. R.↗

Domain decomposition methods for the parallel computation of reacting flows

Domain decomposition is a natural route to parallel computing for partial differential equation solvers. Subdomains of which the original domain of definition is comprised are assigned to independent processors at the price of periodic coordination between processors to compute global parameters and maintain the requisite degree of continuity of the solution at the subdomain interfaces. In the domain-decomposed solution of steady multidimensional systems of PDEs by finite difference methods using a pseudo-transient version of Newton iteration, the only portion of the computation which generally stands in the way of efficient parallelization is the solution of the large, sparse linear systems arising at each Newton step. For some Jacobian matrices drawn from an actual two-dimensional reacting flow problem, comparisons are made between relaxation-based linear solvers and also preconditioned iterative methods of Conjugate Gradient and Chebyshev type, focusing attention on both iteration count and global inner product count. The generalized minimum residual method with block-ILU preconditioning is judged the best serial method among those considered, and parallel numerical experiments on the Encore Multimax demonstrate for it approximately 10-fold speedup on 16 processors.

Keyes, David E.↗

Neural controller for adaptive movements with unforeseen payloads

A theory and computer simulation of a neural controller that learns to move and position a link carrying an unforeseen payload accurately are presented. The neural controller learns adaptive dynamic control from its own experience. It does not use information about link mass, link length, or direction of gravity, and it uses only indirect uncalibrated information about payload and actuator limits. Its average positioning accuracy across a large range of payloads after learning is 3 percent of the positioning range. This neural controller can be used as a basis for coordinating any number of sensory inputs with limbs of any number of joints. The feedforward nature of control allows parallel implementation in real time across multiple joints.

Kuperstein, Michael↗

Dispersion relation for bianisotropic materials and its symmetry properties

The dispersion relation for an arbitrary general bianisotropic medium is derived in Cartesian coordinates, in a form well suited to imposing the boundary conditions when dealing with layered media with planar and parallel interfaces. Special cases of practical interest are also considered. Eleven fundamental coefficient families are identified by considering in detail all the symmetries present in the dispersion relation. An ad hoc expression of the determinant of the sum of two 3 x 3 matrices permits the use of a simple procedure to obtain the coefficients of the dispersion equation. The discussed symmetry properties have general validity, and this technique to evaluate the coefficients may be useful in other fields of application where dispersion relations are of importance.

Graglia, Roberto D.↗

Ignition of confined gaseous mixtures by hot surfaces and hot wires

Ignition times and spatial and temporal variations of temperature and concentration in gaseous mixtures confined between two infinite parallel walls or two infinite cylinders have been obtained by numerical integration of the appropriate conservation equations written in Lagrangian coordinates. Ignition times and ignition energies are presented for the case of an isothermal wall in terms of the initial mixture pressure and equivalence ratio for both one and two-step chemical reaction mechanisms. The numerical results indicate that there is a critical mixture pressure for which the ignition time is minimum. The values of this critical pressure are larger (smaller) than 1 atm for the one- (two-) step reaction mechanism. The critical pressure for the ignition time is not equal to the critical pressure for the ignition energy. The ignition time and energy decrease with the equivalence ratio within a certain range and then remain constant.

Ramos, J. I.↗

Supercomputing 2002: NAS Demo Abstracts

The hyperwall is a new concept in visual supercomputing, conceived and developed by the NAS Exploratory Computing Group. The hyperwall will allow simultaneous and coordinated visualization and interaction of an array of processes, such as a the computations of a parameter study or the parallel evolutions of a genetic algorithm population. Making over 65 million pixels available to the user, the hyperwall will enable and elicit qualitatively new ways of leveraging computers to accomplish science. It is currently still unclear whether we will be able to transport the hyperwall to SC02. The crucial display frame still has not been completed by the metal fabrication shop, although they promised an August delivery. Also, we are still working the fragile node issue, which may require transplantation of the compute nodes from the present 2U cases into 3U cases. This modification will increase the present 3-rack configuration to 5 racks.

Parks, John↗

Software for Verifying Image-Correlation Tie Points

A computer program enables assessment of the quality of tie points in the image-correlation processes of the software described in the immediately preceding article. Tie points are computed in mappings between corresponding pixels in the left and right images of a stereoscopic pair. The mappings are sometimes not perfect because image data can be noisy and parallax can cause some points to appear in one image but not the other. The present computer program relies on the availability of a left- right correlation map in addition to the usual right left correlation map. The additional map must be generated, which doubles the processing time. Such increased time can now be afforded in the data-processing pipeline, since the time for map generation is now reduced from about 60 to 3 minutes by the parallelization discussed in the previous article. Parallel cluster processing time, therefore, enabled this better science result. The first mapping is typically from a point (denoted by coordinates x,y) in the left image to a point (x',y') in the right image. The second mapping is from (x',y') in the right image to some point (x",y") in the left image. If (x,y) and(x",y") are identical, then the mapping is considered perfect. The perfect-match criterion can be relaxed by introducing an error window that admits of round-off error and a small amount of noise. The mapping procedure can be repeated until all points in each image not connected to points in the other image are eliminated, so that what remains are verified correlation data.

Klimeck, Gerhard↗

Analysis of impingement heat transfer for two parallel liquid-metal slot jets

An analytical method is developed for determining heat transfer by impinging liquid-metal slot jets. The method involves mapping the jet flow region, which is bounded by free streamlines, into a potential plane where it becomes a uniform flow in a channel of constant width. The energy equation is transformed into potential plane coordinates and is solved in the channel flow region. Conformal mapping is then used to transform the solution back into the physical plane and obtain the desired heat-transfer characteristics. The analysis given here determines the heat-transfer characteristics for two parallel liquid-metal slot jets impinging normally against a uniformly heated flat plate. The liquid-metal assumptions are made that the jets are inviscid and that molecular conduction is dominating heat diffusion. Wall temperature distributions along the heated plate are obtained as a function of spacing between the jets and the jet Peclet number.

Siegel, R.↗

Implementation of a Parallel Kalman Filter for Stratospheric Chemical Tracer Assimilation

A Kalman filter for the assimilation of long-lived atmospheric chemical constituents has been developed for two-dimensional transport models on isentropic surfaces over the globe. An important attribute of the Kalman filter is that it calculates error covariances of the constituent fields using the tracer dynamics. Consequently, the current Kalman-filter assimilation is a five-dimensional problem (coordinates of two points and time), and it can only be handled on computers with large memory and high floating point speed. In this paper, an implementation of the Kalman filter for distributed-memory, message-passing parallel computers is discussed. Two approaches were studied: an operator decomposition and a covariance decomposition. The latter was found to be more scalable than the former, and it possesses the property that the dynamical model does not need to be parallelized, which is of considerable practical advantage. This code is currently used to assimilate constituent data retrieved by limb sounders on the Upper Atmosphere Research Satellite. Tests of the code examined the variance transport and observability properties. Aspects of the parallel implementation, some timing results, and a brief discussion of the physical results will be presented.

Chang, Lang-Ping↗