Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel codes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25

Methodology of modeling and measuring computer architectures for plasma simulations

A brief introduction to plasma simulation using computers and the difficulties on currently available computers is given. Through the use of an analyzing and measuring methodology - SARA, the control flow and data flow of a particle simulation model REM2-1/2D are exemplified. After recursive refinements the total execution time may be greatly shortened and a fully parallel data flow can be obtained. From this data flow, a matched computer architecture or organization could be configured to achieve the computation bound of an application problem. A sequential type simulation model, an array/pipeline type simulation model, and a fully parallel simulation model of a code REM2-1/2D are proposed and analyzed. This methodology can be applied to other application problems which have implicitly parallel nature.

Wang, L. P. T.↗

Converting Time Signals From BCD to IRIG-B

Coded representation of time signals--day, hour, minute, second--is changed from binary-coded decimal (BCD) to IRIG standard time-code format B by circuit that uses nine integrated circuits. Input to code-converter circuit is parallel BCD pulses on bus output is serial pulses of IRIG-B on single line.

Houston, J. B.↗

Increasing processor utilization during parallel computation rundown

Some parallel processing environments provide for asynchronous execution and completion of general purpose parallel computations from a single computational phase. When all the computations from such a phase are complete, a new parallel computational phase is begun. Depending upon the granularity of the parallel computations to be performed, there may be a shortage of available work as a particular computational phase draws to a close (computational rundown). This can result in the waste of computing resources and the delay of the overall problem. In many practical instances, strict sequential ordering of phases of parallel computation is not totally required. In such cases, the beginning of one phase can be correctly computed before the end of a previous phase is completed. This allows additional work to be generated somewhat earlier to keep computing resources busy during each computational rundown. The conditions under which this can occur are identified and the frequency of occurrence of such overlapping in an actual parallel Navier-Stokes code is reported. A language construct is suggested and possible control strategies for the management of such computational phase overlapping are discussed.

Jones, W. H.↗

Phase space simulation of collisionless stellar systems on the massively parallel processor

A numerical technique for solving the collisionless Boltzmann equation describing the time evolution of a self gravitating fluid in phase space was implemented on the Massively Parallel Processor (MPP). The code performs calculations for a two dimensional phase space grid (with one space and one velocity dimension). Some results from calculations are presented. The execution speed of the code is comparable to the speed of a single processor of a Cray-XMP. Advantages and disadvantages of the MPP architecture for this type of problem are discussed. The nearest neighbor connectivity of the MPP array does not pose a significant obstacle. Future MPP-like machines should have much more local memory and easier access to staging memory and disks in order to be effective for this type of problem.

White, Richard L.↗

Parallel software support for computational structural mechanics

The application of the parallel programming methodology known as the Force was conducted. Two application issues were addressed. The first involves the efficiency of the implementation and its completeness in terms of satisfying the needs of other researchers implementing parallel algorithms. Support for, and interaction with, other Computational Structural Mechanics (CSM) researchers using the Force was the main issue, but some independent investigation of the Barrier construct, which is extremely important to overall performance, was also undertaken. Another efficiency issue which was addressed was that of relaxing the strong synchronization condition imposed on the self-scheduled parallel DO loop. The Force was extended by the addition of logical conditions to the cases of a parallel case construct and by the inclusion of a self-scheduled version of this construct. The second issue involved applying the Force to the parallelization of finite element codes such as those found in the NICE/SPAR testbed system. One of the more difficult problems encountered is the determination of what information in COMMON blocks is actually used outside of a subroutine and when a subroutine uses a COMMON block merely as scratch storage for internal temporary results.

Jordan, Harry F.↗

Increasing processor utilization during parallel computation rundown

Some parallel processing environments provide for asynchronous execution and completion of general purpose parallel computations from a single computational phase. When all the computations from such a phase are complete, a new parallel computational phase is begun. Depending upon the granularity of the parallel computations to be performed, there may be a shortage of available work as a particular computational phase draws to a close (computational rundown). This can result in the waste of computing resources and the delay of the overall problem. In many practical instances, strict sequential ordering of phases of parallel computation is not totally required. In such cases, the beginning of one phase can be correctly computed before the end of a previous phase is completed. This allows additional work to be generated somewhat earlier to keep computing resources busy during each computational rundown. The conditions under which this can occur are identified and the frequency of occurrence of such overlapping in an actual parallel Navier-Stokes code is reported. A language construct is suggested and possible control strategies for the management of such computational phase overlapping are discsused.

Jones, William H.↗

Analysis and simulations of a troposphere-stratosphere gravity wave model. I

An analytical model is presented that accommodates nonhydrostatic shearing stratified flow over an obstacle, and that can be modified to include a superposed stratosphere with constant wind and higher stability. A simulation code is used in parallel with the analytic calculations to demonstrate a methodology for determining the wavelength and magnitude of gravity wave energy reflected, and that transmitted, by the tropopause.

Wurtele, M. G.↗

Voyager image data compression and block encoding

Telemetry enhancement techniques used by Voyager-2 to reduce telemetry transmission rates by over 50 percent compared to those used at Saturn, with negligible loss in information return, are described. The use of the Reed-Solomon encoder is discussed, and the principles and implementation of an Image Data Compressor algorithm for noiseless coding techniques are addressed. Parallel operation of the redundant Flight Data Subsystem processors is discussed.

Urban, Michael G.↗

Programming Probabilistic Structural Analysis for Parallel Processing Computer

The ultimate goal of this research program is to make Probabilistic Structural Analysis (PSA) computationally efficient and hence practical for the design environment by achieving large scale parallelism. The paper identifies the multiple levels of parallelism in PSA, identifies methodologies for exploiting this parallelism, describes the development of a parallel stochastic finite element code, and presents results of two example applications. It is demonstrated that speeds within five percent of those theoretically possible can be achieved. A special-purpose numerical technique, the stochastic preconditioned conjugate gradient method, is also presented and demonstrated to be extremely efficient for certain classes of PSA problems.

Sues, Robert H.↗

Seal development activities at Allison Turbine Division

Brush seals are being evaluated for potential near and far term gas turbine engine applications. Development is in the form of rig component testing and engine testing. Allison has tested an engine with 20 individual brush seal positions. These seals were located throughout the engine. The emphasis of the current work is on obtaining long term performance data for brush seals. Very little of this data is available. Allison is presently developing film riding face seal technology to support future gas turbine engine applications. A face seal with an approximate 7 inch diameter was successfully tested to 1000 F, 100 psid, and 650 ft/sec. Seal leakage remained below 1 scfm throughout the duration of the test. A model for the compressible gas film was developed which separates the model for the compressible gas film was developed which separates the primary seal rings during operation. This model is based on the traditional Reynold's approach which is customarily applied to lubrication type problems. Because of the difficulty of experimentally verifying the program predictions, a commercial Navier-Stokes code was used in parallel. By comparing predictions for similar cases, it is expected that the limitations of the Reynold's model can be assessed as it applies to this particular seal.

Munson, John↗

A portable MPI-based parallel vector template library

This paper discusses the design and implementation of a polymorphic collection library for distributed address-space parallel computers. The library provides a data-parallel programming model for C++ by providing three main components: a single generic collection class, generic algorithms over collections, and generic algebraic combining functions. Collection elements are the fourth component of a program written using the library and may be either of the built-in types of C or of user-defined types. Many ideas are borrowed from the Standard Template Library (STL) of C++, although a restricted programming model is proposed because of the distributed address-space memory model assumed. Whereas the STL provides standard collections and implementations of algorithms for uniprocessors, this paper advocates standardizing interfaces that may be customized for different parallel computers. Just as the STL attempts to increase programmer productivity through code reuse, a similar standard for parallel computers could provide programmers with a standard set of algorithms portable across many different architectures. The efficacy of this approach is verified by examining performance data collected from an initial implementation of the library running on an IBM SP-2 and an Intel Paragon.

Sheffler, Thomas J.↗

A Portable MPI-Based Parallel Vector Template Library

This paper discusses the design and implementation of a polymorphic collection library for distributed address-space parallel computers. The library provides a data-parallel programming model for C + + by providing three main components: a single generic collection class, generic algorithms over collections, and generic algebraic combining functions. Collection elements are the fourth component of a program written using the library and may be either of the built-in types of c or of user-defined types. Many ideas are borrowed from the Standard Template Library (STL) of C++, although a restricted programming model is proposed because of the distributed address-space memory model assumed. Whereas the STL provides standard collections and implementations of algorithms for uniprocessors, this paper advocates standardizing interfaces that may be customized for different parallel computers. Just as the STL attempts to increase programmer productivity through code reuse, a similar standard for parallel computers could provide programmers with a standard set of algorithms portable across many different architectures. The efficacy of this approach is verified by examining performance data collected from an initial implementation of the library running on an IBM SP-2 and an Intel Paragon.

Sheffler, Thomas J.↗

Multistage Simulations of the GE90 Turbine

The average passage approach has been used to analyze three multistage configurations of the GE90 turbine. These are a high pressure turbine rig, a low pressure turbine rig and a full turbine configuration comprising 18 blade rows of the GE90 engine at takeoff conditions. Cooling flows in the high pressure turbine have been simulated using source terms. This is the first time a dual-spool cooled turbine has been analyzed in 3D using a multistage approach. There is good agreement between the simulations and experimental results. Multistage and component interaction effects are also presented. The parallel efficiency of the code is excellent at 87.3% using 121 processors on an SGI Origin for the 18 blade row configuration. The accuracy and efficiency of the calculation now allow it to be effectively used in a design environment so that multistage effects can be accounted for in turbine design.

Turner, Mark G.↗

Implementation of an Eta Belt Domain on Parallel Systems

We extend the Eta weather model from a regional domain into a belt domain that does not require meridional boundary conditions. We describe how the extension is achieved and the parallel implementation of the code on the Cray T3E and the SGI Origin 2000. We validate the forecast results on the two platforms and examine how the removal of the meridional boundary conditions affects these forecasts. In addition, using several domains of different sizes and resolutions, we present the scaling performance of the code on both systems.

Kouatchou, Jules↗

Charon Toolkit for Parallel, Implicit Structured-Grid Computations: Functional Design

Charon is a software toolkit that enables engineers to develop high-performing message-passing programs in a convenient and piecemeal fashion. Emphasis is on rapid program development and prototyping. In this report a detailed description of the functional design of the toolkit is presented. It is illustrated by the stepwise parallelization of two representative code examples.

VanderWijngaart, Rob F.↗

Collaborative Simulation Grid: Multiscale Quantum-Mechanical/Classical Atomistic Simulations on Distributed PC Clusters in the US and Japan

A multidisciplinary, collaborative simulation has been performed on a Grid of geographically distributed PC clusters. The multiscale simulation approach seamlessly combines i) atomistic simulation backed on the molecular dynamics (MD) method and ii) quantum mechanical (QM) calculation based on the density functional theory (DFT), so that accurate but less scalable computations are performed only where they are needed. The multiscale MD/QM simulation code has been Grid-enabled using i) a modular, additive hybridization scheme, ii) multiple QM clustering, and iii) computation/communication overlapping. The Gridified MD/QM simulation code has been used to study environmental effects of water molecules on fracture in silicon. A preliminary run of the code has achieved a parallel efficiency of 94% on 25 PCs distributed over 3 PC clusters in the US and Japan, and a larger test involving 154 processors on 5 distributed PC clusters is in progress.

Kikuchi, Hideaki↗

High Performance Parallel Methods for Space Weather Simulations

This is the final report of our NASA AISRP grant entitled 'High Performance Parallel Methods for Space Weather Simulations'. The main thrust of the proposal was to achieve significant progress towards new high-performance methods which would greatly accelerate global MHD simulations and eventually make it possible to develop first-principles based space weather simulations which run much faster than real time. We are pleased to report that with the help of this award we made major progress in this direction and developed the first parallel implicit global MHD code with adaptive mesh refinement. The main limitation of all earlier global space physics MHD codes was the explicit time stepping algorithm. Explicit time steps are limited by the Courant-Friedrichs-Lewy (CFL) condition, which essentially ensures that no information travels more than a cell size during a time step. This condition represents a non-linear penalty for highly resolved calculations, since finer grid resolution (and consequently smaller computational cells) not only results in more computational cells, but also in smaller time steps.

Hunter, Paul↗

Numerical Simulation of Protoplanetary Vortices

The fluid dynamics within a protoplanetary disk has been attracting the attention of many researchers for a few decades. Previous works include, to list only a few among many others, the well-known prescription of Shakura & Sunyaev, the convective and instability study of Stone & Balbus and Hawley et al., the Rossby wave approach of Lovelace et al., as well as a recent work by Klahr & Bodenheimer, which attempted to identify turbulent flow within the disk. The disk is commonly understood to be a thin gas disk rotating around a central star with differential rotation (the Keplerian velocity), and the central quest remains as how the flow behavior deviates (albeit by a small amount) from a strong balance established between gravitational and centrifugal forces, transfers mass and momentum inward, and eventually forms planetesimals and planets. In earlier works we have briefly described the possible physical processes involved in the disk; we have proposed the existence of long-lasting, coherent vortices as an efficient agent for mass and momentum transport. In particular, Barranco et al. provided a general mathematical framework that is suitable for the asymptotic regime of the disk; Barranco & Marcus (2000) addressed a proposed vortex-dust interaction mechanism which might lead to planetesimal formation; and Lin et al. (2002), as inspired by general geophysical vortex dynamics, proposed basic mechanisms by which vortices can transport mass and angular momentum. The current work follows up on our previous effort. We shall focus on the detailed numerical implementation of our problem. We have developed a parallel, pseudo-spectral code to simulate the full three-dimensional vortex dynamics in a stably-stratified, differentially rotating frame, which represents the environment of the disk. Our simulation is validated with full diagnostics and comparisons, and we present our results on a family of three-dimensional, coherent equilibrium vortices.

Lin, H.↗