Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel codes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26

Magnetospheric Magnetic Reconnection with Southward IMF by a 3D EMPM Simulation

We report our new simulation results on magnetospheric magnetic reconnection with southward IMF using a 3D EMPM model, with greater resolution and more particles using the parallelized 3D HPF TRISTAN code on VPP5000 supercomputer. Main parameters used in the new simulation are: domain size is 215 x 145 x 145, grid size = 0.5 Earth radius, initial particle number is 16 per cell, the IMF is southward. Arrival of southward IMF will cause reconnection in the magnetopause, thus allowing particles to enter into the inner magnetosphere. Sunward and tailward high particle flow are observed by satellites, and these phenomena are also observed in the simulation near the neutral line (X line) of the near-Earth magnetotail. This high particle flow goes along with the reconnected island. The magnetic reconnection process contributes to direct plasma entry between the magnetosheath to the inner magnetosphere and plasma sheet, in which the entry process eats the magnetosheath plasma to plasma sheet temperatures. We investigate magnetic, electric fields, density, and current during this magnetic reconnection with southward IMF. Further investigation with this simulation will provide insight into unsolved problems, such as the triggering of storms and substorms, and the storm-substorm relationship. New results will be presented at the meeting.

Nishikawa, K.-I.↗

Scalable Implementation of Finite Elements by NASA _ Implicit (ScIFEi)

Scalable Implementation of Finite Elements by NASA (ScIFEN) is a parallel finite element analysis code written in C++. ScIFEN is designed to provide scalable solutions to computational mechanics problems. It supports a variety of finite element types, nonlinear material models, and boundary conditions. This report provides an overview of ScIFEi (\Sci-Fi"), the implicit solid mechanics driver within ScIFEN. A description of ScIFEi's capabilities is provided, including an overview of the tools and features that accompany the software as well as a description of the input and output le formats. Results from several problems are included, demonstrating the efficiency and scalability of ScIFEi by comparing to finite element analysis using a commercial code.

Warner, James E.↗

Progress of High Efficiency Centrifugal Compressor Simulations Using TURBO

Three-dimensional, time-accurate, and phase-lagged computational fluid dynamics (CFD) simulations of the High Efficiency Centrifugal Compressor (HECC) stage were generated using the TURBO solver. Changes to the TURBO Parallel Version 4 source code were made in order to properly model the no-slip boundary condition along the spinning hub region for centrifugal impellers. A startup procedure was developed to generate a converged flow field in TURBO. This procedure initialized computations on a coarsened mesh generated by the Turbomachinery Gridding System (TGS) and relied on a method of systematically increasing wheel speed and backpressure. Baseline design-speed TURBO results generally overpredicted total pressure ratio, adiabatic efficiency, and the choking flow rate of the HECC stage as compared with the design-intent CFD results of Code Leo. Including diffuser fillet geometry in the TURBO computation resulted in a 0.6 percent reduction in the choking flow rate and led to a better match with design-intent CFD. Diffuser fillets reduced annulus cross-sectional area but also reduced corner separation, and thus blockage, in the diffuser passage. It was found that the TURBO computations are somewhat insensitive to inlet total pressure changing from the TURBO default inlet pressure of 14.7 pounds per square inch (101.35 kilopascals) down to 11.0 pounds per square inch (75.83 kilopascals), the inlet pressure of the component test. Off-design tip clearance was modeled in TURBO in two computations: one in which the blade tip geometry was trimmed by 12 mils (0.3048 millimeters), and another in which the hub flow path was moved to reflect a 12-mil axial shift in the impeller hub, creating a step at the hub. The one-dimensional results of these two computations indicate non-negligible differences between the two modeling approaches.

turbomachinery↗

ANALYSIS OF THE MSL/MEDLI ENTRY DATA WITH COUPLED CFD AND MATERIAL RESPONSE.

The Mars Science Laboratory (MSL) was protected during its atmospheric entry by an instrumented heat-shield using NASA's Phenolic Impregnated Carbon Ablator (PICA) material. PICA is a lightweight carbon fiber/polymeric resin material that offers out-standing performances for protecting probes during planetary entry. The Mars Entry Descent and Landing Instrument (MEDLI) suite on MSL offers unique in-flight validation data for models of material response and atmospheric entry. MEDLI recorded, among other things, time-resolved in-depth temperature data of PICA using thermocouple sensors assembled in the MEDLI Integrated Sensor Plugs (MISP). The objective of this work is to showcase and analyze the coupling between the material response and the aerothermal environment. As shown in Figure 1, the workflow is divided into the following steps. First, the aerothermal properties are computed in the Data Parallel Line Relaxation (DPLR) code [3] and used with the Nonequilibrium air radiation (NEQAIR) program [8] to compute radiative heating. Second, the thermal response inside the material is computed in the Porous material Analysis Toolbox based on Open-FOAM (PATO) using a fixed blowing correction parameter. Third, the pyrolysis gases computed in PATO are used as inputs to a blowing boundary condition within DPLR. Fourth, the new environment properties from DPLR are used in NEQAIR to provide an updated solution, then both the updated aerothermal environment and radiative heating are used in PATO without blowing correction. The third and fourth steps are then repeated until convergence in surface temperature is obtained. Convergence in the radiative heating is generally achieved before surface temperature, at which point the radiative heating is no longer updated. Char mass loss rates are forced to zero to produce a non-receding surface condition. For early time points in the trajectory, where flow around the MSL aeroshell is rarefied, the Direct Simulation Monte Carlo (DSMC) code, SPARTA, is used to compute the aerothermal environment. Iteration between PATO and SPARTA is not performed due to the computational cost of DSMC simulations. Preliminary results of the coupling between PATO and DPLR for the MSL heatshield atmospheric entry model are presented in Figures 2-4 at 65 seconds after entry interface. Figure 2 shows the surface temperature results from an uncoupled simulation in PATO with the blowing correction parameter applied (left) along with the coupled surface temperature after iteration (right). Figure 3 shows the surface temperature along the centerline from windward to leeward for easier comparison. Figure 4 shows the coupled and uncoupled pyrolysis gas blowing rate. Mars 2020 used a similar heatshield consisting of PICA for thermal protection during entry, descent, and landing. In preparation for Mars 2020 post-flight analysis, the predictive material response capability is benchmarked against flight data from MEDLI. This work represents an important milestone toward the development of validated predictive capabilities for designing thermal protection systems for planetary probes.

Mars Science Laboratory↗

Analysis of MSL/MEDLI Entry Data with Coupled CFD and Material Response

The Mars Science Laboratory (MSL) was protected during its atmospheric entry by an instrumented heatshield using NASA's Phenolic Impregnated Carbon Ablator (PICA) material [1]. PICA is a lightweight carbon fiber/polymeric resin material that offers outstanding performances for protecting probes during planetary entry. The Mars Entry Descent and Landing Instrument (MEDLI) suite on MSL offers unique in-flight validation data for models of material response and atmospheric entry. MEDLI recorded, among other things, time-resolved in-depth temperature data of PICA using thermocouple sensors assembled in the MEDLI Integrated Sensor Plugs (MISP) [2]. The objective of this work is to showcase and analyze the coupling between the material response and the aerothermal environment. As shown in Figure 1, the workflow is divided into the following steps. First, the aerothermal properties are computed in the Data Parallel Line Relaxation (DPLR) code [3] and used with the Nonequilibrium air radiation (NEQAIR) program [8] to compute radiative heating. Second, the thermal response inside the material is computed in the Porous material Analysis Toolbox based on OpenFOAM (PATO) [4,5,6] using a fixed blowing correction parameter. Third, the pyrolysis gases computed in PATO are used as inputs to a blowing boundary condition within DPLR. Fourth, the new environment properties from DPLR are used in NEQAIR to provide an updated solution, then both the updated aerothermal environment and radiative heating are used in PATO without blowing correction. The third and fourth steps are then repeated until convergence in surface temperature is obtained. Convergence in the radiative heating is generally achieved before surface temperature, at which point the radiative heating is no longer updated. Char mass loss rates are forced to zero to produce a non-receding surface condition. For early time points in the trajectory, where flow around the MSL aeroshell is rarefied, the Direct Simulation Monte Carlo (DSMC) code, SPARTA [7], is used to compute the aerothermal environment. Iteration between PATO and SPARTA is not performed due to the computational cost of DSMC simulations. Preliminary results of the coupling between PATO and DPLR for the MSL heatshield atmospheric entry model are presented in Figures 2-4 at 65 seconds after entry interface. Figure 2 shows the surface temperature results from an uncoupled simulation in PATO with the blowing correction parameter applied (left) along with the coupled surface temperature after iteration (right). Figure 3 shows the surface temperature along the centerline from windward to leeward for easier comparison. Figure 4 shows the coupled and uncoupled pyrolysis gas blowing rate. Mars 2020 used a similar heatshield consisting of PICA for thermal protection during entry, descent, and landing. In preparation for Mars 2020 post-flight analysis, the predictive material response capability is benchmarked against flight data from MEDLI. This work represents an important milestone toward the development of validated predictive capabilities for designing thermal protection systems for planetary probes.

Thermal Protection Systems↗

NASA Langley FUN3D Analyses in Support of the 1st AIAA Stability and Control Prediction Workshop

This work summarizes the results of FUN3D analyses conducted for the 1st AIAA Stability and Control Workshop on behalf of participants from the NASA Langley Research Center. The workshop was created to establish best practices for the prediction of stability and control derivatives using computational fluid dynamics and assess the limitations of these methods when those best practices are applied. The inaugural workshop considered the ONERA version of the NASA/Boeing Common Research Model, which includes the wing, body, horizontal tail, and a vertical tail designed by ONERA. Wind tunnel data at small sideslip angles remain unpublished and served as ‘blind’ data for computational comparisons. The present research generated workshop test case data using the NASA FUN3D code, which is a parallelized, unstructured, node-based, finite-volume discretization, Reynolds-averaged Navier-Stokes flow solver. Steady- state numerical simulations were conducted for workshop test cases investigating the following: grid convergence, Mach number effect on static stability, wind tunnel sting increments, static stability-derivative calculations, and a sideslip angle sweep. Results were generated for two series of unstructured, mixed-element grids, one set provided by the workshop and another set created using the HeldenMesh grid generation software. The results provided include total- and component-level breakdowns of the force and moment coefficients, in addition to sectional pressure distributions for the wing and tail components for comparisons to wind tunnel data.

CFD↗

Identification and Study of Validation Level Test Cases for Computational Modeling of Non-Charring Ablators

Computational modeling of Thermal Protection System (TPS) materials, used for aerospace applications, provides numerous advantages in preliminary selection and design of a heatshield material and shape for atmospheric entry vehicles. However, to serve as a reliable tool for prediction of material thermal and ablative behavior, the modeling approach needs to be validated against real experimental and flight data, preferably at a range of applied conditions. The validation study is typically very complex as it requires reliable measured data not only for the material thermal response and surface recession, but also well characterized environmental conditions. The validation problem becomes even more complex when the material thermal response is dictated by multi-physics effects such as solid conduction, in-depth thermal decomposition, pyrolysis gas flow and chemical reactions. The multi-physics effects complicate not only the modeling effort, but also the experimental measurement for validation of various aspects of the highly coupled problem. In this study, an attempt is made to identify suitable experimental data that could serve as a source for validation of material thermal response modeling tools. To reduce the computational complexity, this study focuses only on non-charring ablators, where the material thermal response could be modeled with a single governing equation for solid conduction and the ablation is limited only to the surface of the material. With a well characterized and publicly available experimental data being sparse, the study is limited in presenting test cases for only three materials: camphor, graphite and FiberForm® in the sequence of increased modeling complexity. Graphite is a commonly used TPS material for aerospace applications, both for leading edges of high-speed vehicles and internal insulation of solid rocket motors. FiberForm® is a porous carbon pre-form used in preparation of the well known PICA material Tran et al. [1996]. Inclusion of camphor into the list is conditioned with the relative simplicity in modeling the material thermal and chemical response and the low-enthalpy flow environment. In addition, camphor has been used as a simple test material for study of flow transition behavior by Stock and Ginoux [1973] and assessment of a heatshield shape change at flight relevant conditions by Rotondi et al. [2022]. In this work, the identified experimental data was extracted from the public literature and test cases that yet have been published. As it was found from the review, not a single test case contains an exhaustive set of data that would validate every aspect of the material physics. However, in the data collected, various aspects of the material behavior can be still validated, such as surface and in-depth temperature, amount of recession and a shape change. The identified experimental data for each case is accompanied with a characterized flow environment and simulated boundary conditions predicted by a Data-Parallel Line Relaxation (DPLR) code Wright et al. [1998]. In addition, material thermal response numerical simulations in each test case are performed with Kentucky Aerothermodynamics and Thermal Response System (KATS-MR) Zibitsker et al. [2022] providing a comparative study and a sanity check for the proposed validation data. Sample results from the performed numerical study are shown below. Figure 1 shows distribution of surface heat flux and pressure values on a hemi-cylinder model made of FiberForm® and tested in HyMETS arc-jet facility. The results are shown for the high pressure condition among the two tests. Flow simulation was performed with DPLR code on a quarter of original geometry. In the figure, the quarter shape was mirrored across zx and xy planes to show the complete distribution. Figure 2 shows the material response results for the high pressure case (7500 Pa), simulated with KATS-MR and a comparison to the experimental data for the surface temperature and shape shape. The simulation was performed on a 2-D slice, extracted in the xy plane at the middle of the sample. Figure 3 shows the material response simulation for the low pressure case (3500 Pa) and a comparison to the experimental data for surface temperature and shape change.

ablation↗

Parallel Aeroelastic Analysis Using ENSAERO and NASTRAN

A high fidelity parallel static structural analysis capability is created and interfaced to the multidisciplinary analysis package ENSAERO-MPI of Ames Research Center. This new module replaced ENSAERO's lower fidelity simple finite element and modal modules. Full aircraft structures may be more accurately modeled using the new finite element capability. Parallel computation is performed by breaking the full structure into multiple substructures. This approach is conceptually similar to ENSAERO's multi-zonal fluid analysis capability. The new substructure code is used to solve the structural finite element equations for each substructure in parallel. NASTRAN/COSMIC is utilized as a front end for this code. Its full library of elements can be used to create an accurate and realistic aircraft mode. It is used to create the stiffness matrices for each sub-structure. The new parallel code then uses an iterative preconditioned conjugate gradient method to solve the global structural equations for the sub-structure boundary nodes. Results are presented for a wing-body configuration.

Eldred, Lloyd B.↗

High speed two-dimensional event detection and imaging using an analog interface and a massively parallel processor

A quantitative pulse count (event detection) algorithm with linearity to high count rates is accomplished by combining a high-speed, high frame rate camera with simple logic code run on a massively parallel processor such as a GPU or a FPGA. The parallel processor elements examine frames from the camera pixel by pixel to find and tag events or count pulses. The tagged events are combined to form a combined quantitative event image.

Waugh, Justin↗

Explicit structure-preserving geometric particle-in-cell algorithm in curvilinear orthogonal coordinate systems and its applications to whole-device 6D kinetic simulations of tokamak physics

Explicit structure-preserving geometric particle-in-cell (PIC) algorithm in curvilinear orthogonal coordinate systems is developed. The work reported represents a further development of the structure-preserving geometric PIC algorithm achieving the goal of practical applications in magnetic fusion research. The algorithm is constructed by discretizing the field theory for the system of charged particles and electromagnetic field using Whitney forms, discrete exterior calculus, and explicit non-canonical symplectic integration. In addition to the truncated infinitely dimensional symplectic structure, the algorithm preserves exactly many important physical symmetries and conservation laws, such as local energy conservation, gauge symmetry and the corresponding local charge conservation. As a result, the algorithm possesses the long-term accuracy and fidelity required for first-principles-based simulations of the multiscale tokamak physics. The algorithm has been implemented in the SymPIC code, which is designed for high-efficiency massively-parallel PIC simulations in modern clusters. The code has been applied to carry out whole-device 6D kinetic simulation studies of tokamak physics. A self-consistent kinetic steady state for fusion plasma in the tokamak geometry is numerically found with a predominately diagonal and anisotropic pressure tensor. The state also admits a steady-state sub-sonic ion flow in the range of 10 km s -1 , agreeing with experimental observations and analytical calculations Kinetic ballooning instability in the self-consistent kinetic steady state is simulated. It is shown that high-n ballooning modes have larger growth rates than low-n global modes, and in the nonlinear phase the modes saturate approximately in 5 ion transit times at the 2% level by the E × B flow generated by the instability. These results are consistent with early and recent electromagnetic gyrokinetic simulations.

43 PARTICLE ACCELERATORS↗

ADPAC v1.0: User's Manual

The overall objective of this study was to evaluate the effects of turbulence models in a 3-D numerical analysis on the wake prediction capability. The current version of the computer code resulting from this study is referred to as ADPAC v7 (Advanced Ducted Propfan Analysis Codes -Version 7). This report is intended to serve as a computer program user's manual for the ADPAC code used and modified under Task 15 of NASA Contract NAS3-27394. The ADPAC program is based on a flexible multiple-block and discretization scheme permitting coupled 2-D/3-D mesh block solutions with application to a wide variety of geometries. Aerodynamic calculations are based on a four-stage Runge-Kutta time-marching finite volume solution technique with added numerical dissipation. Steady flow predictions are accelerated by a multigrid procedure. Turbulence models now available in the ADPAC code are: a simple mixing-length model, the algebraic Baldwin-Lomax model with user defined coefficients, the one-equation Spalart-Allmaras model, and a two-equation k-R model. The consolidated ADPAC code is capable of executing in either a serial or parallel computing mode from a single source code.

Hall, Edward J.↗

Parallel computation in a three-dimensional elastic-plastic finite-element analysis

CRAY 2 and CRAY Y-MP evaluation runs are undertaken for a three-dimensional, elastoplastic FEM code implementation of the 'autotasking' parallel-processing technique, which is used in all code components except the matrix-equation solver. For a typical example problem in the case of a four-processor CRAY 2, a speedup factor of 2.1 was found achievable in a dedicated environment, and 1.7 in a multiuser environment; for a three-processor CRAY Y-MP, speedup in a multiuser environment was about 2.4. Only minimal effort was required to implement autotasking in this implementation code.

Shivakumar, K. N.↗

Enabling Parallel Execution of System-level Simulations in SAM

This report summarizes the recent code updates related to “element ghosting” in SAM to enable the parallel execution of system-level simulations using multiple processors/cores. Unlike typical MOOSE-based applications, for system-level simulations, SAM mostly deals with a collection of discrete small pieces of meshes, and the connection of physics on these meshes are realized by using “connector” types of components/code structures, such as conjugate heat transfer and flow junctions. The required code implementation is to correctly mark the necessary ghost elements for each type of such components/code structures; thus, the lower-level libraries can correctly perform the necessary data transfer between processors (CPUs) when executed in parallel mode. After the code updates, SAM can now run system-level simulations in the parallel mode. The parallel execution capability was then tested with an ABTR input model with 23k DOFs. Significant speedup was demonstrated when the optimal number of CPUs were used in parallel mode. Future systematic studies on parallelization performance using additional test cases covering different physics/scenarios will be needed to provide additional insights into the scalability of SAM parallelization.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Good trellises for IC implementation of viterbi decoders for linear block codes

This paper investigates trellis structures of linear block codes for the IC (integrated circuit) implementation of Viterbi decoders capable of achieving high decoding speed while satisfying a constraint on the structural complexity of the trellis in terms of the maximum number of states at any particular depth. Only uniform sectionalizations of the code trellis diagram are considered. An upper bound on the number of parallel and structurally identical (or isomorphic) subtrellises in a proper trellis for a code without exceeding the maximum state complexity of the minimal trellis of the code is first derived. Parallel structures of trellises with various section lengths for binary BCH and Reed-Muller (RM) codes of lengths 32 and 64 are analyzed. Next, the complexity of IC implementation of a Viterbi decoder based on an L-section trellis diagram for a code is investigated. A structural property of a Viterbi decoder called ACS-connectivity which is related to state connectivity is introduced. This parameter affects the complexity of wire-routing (interconnections within the IC). The effect of five parameters namely: (1) effective computational complexity; (2) complexity of the ACS-circuit; (3) traceback complexity; (4) ACS-connectivity; and (5) branch complexity of a trellis diagram on the VLSI complexity of a Viterbi decoder is investigated. It is shown that an IC implementation of a Viterbi decoder based on a non-minimal trellis requires less area and is capable of operation at higher speed than one based on the minimal trellis when the commonly used ACS-array architecture is considered.

Moorthy, H. T.↗

Good trellises for IC implementation of viterbi decoders for linear block codes

This paper investigates trellis structures of linear block codes for the IC (integrated circuit) implementation of Viterbi decoders capable of achieving high decoding speed while satisfying a constraint on the structural complexity of the trellis in terms of the maximum number of states at any particular depth. Only uniform sectionalizations of the code trellis diagram are considered. An upper bound on the number of parallel and structurally identical (or isomorphic) subtrellises in a proper trellis for a code without exceeding the maximum state complexity of the minimal trellis of the code is first derived. Parallel structures of trellises with various section lengths for binary BCH and Reed-Muller (RM) codes of lengths 32 and 64 are analyzed. Next, the complexity of IC implementation of a Viterbi decoder based on an L-section trellis diagram for a code is investigated. A structural property of a Viterbi decoder called ACS-connectivity which is related to state connectivity is introduced. This parameter affects the complexity of wire-routing (interconnections within the IC). The effect of five parameters namely: (1) effective computational complexity; (2) complexity of the ACS-circuit; (3) traceback complexity; (4) ACS-connectivity; and (5) branch complexity of a trellis diagram on the VLSI complexity of a Viterbi decoder is investigated. It is shown that an IC implementation of a Viterbi decoder based on a non-minimal trellis requires less area and is capable of operation at higher speed than one based on the minimal trellis when the commonly used ACS-array architecture is considered.

Lin, Shu↗

Hybrid concatenated codes and iterative decoding

Several improved turbo code apparatuses and methods. The invention encompasses several classes: (1) A data source is applied to two or more encoders with an interleaver between the source and each of the second and subsequent encoders. Each encoder outputs a code element which may be transmitted or stored. A parallel decoder provides the ability to decode the code elements to derive the original source information d without use of a received data signal corresponding to d. The output may be coupled to a multilevel trellis-coded modulator (TCM). (2) A data source d is applied to two or more encoders with an interleaver between the source and each of the second and subsequent encoders. Each of the encoders outputs a code element. In addition, the original data source d is output from the encoder. All of the output elements are coupled to a TCM. (3) At least two data sources are applied to two or more encoders with an interleaver between each source and each of the second and subsequent encoders. The output may be coupled to a TCM. (4) At least two data sources are applied to two or more encoders with at least two interleavers between each source and each of the second and subsequent encoders. (5) At least one data source is applied to one or more serially linked encoders through at least one interleaver. The output may be coupled to a TCM. The invention includes a novel way of terminating a turbo coder.

Divsalar, Dariush↗

The FORCE: A portable parallel programming language supporting computational structural mechanics

This project supports the conversion of codes in Computational Structural Mechanics (CSM) to a parallel form which will efficiently exploit the computational power available from multiprocessors. The work is a part of a comprehensive, FORTRAN-based system to form a basis for a parallel version of the NICE/SPAR combination which will form the CSM Testbed. The software is macro-based and rests on the force methodology developed by the principal investigator in connection with an early scientific multiprocessor. Machine independence is an important characteristic of the system so that retargeting it to the Flex/32, or any other multiprocessor on which NICE/SPAR might be imnplemented, is well supported. The principal investigator has experience in producing parallel software for both full and sparse systems of linear equations using the force macros. Other researchers have used the Force in finite element programs. It has been possible to rapidly develop software which performs at maximum efficiency on a multiprocessor. The inherent machine independence of the system also means that the parallelization will not be limited to a specific multiprocessor.

Jordan, Harry F.↗

NAS Experiences of Porting CM Fortran Codes to HPF on IBM SP2 and SGI Power Challenge

Current Connection Machine (CM) Fortran codes developed for the CM-2 and the CM-5 represent an important class of parallel applications. Several users have employed CM Fortran codes in production mode on the CM-2 and the CM-5 for the last five to six years, constituting a heavy investment in terms of cost and time. With Thinking Machines Corporation's decision to withdraw from the hardware business and with the decommissioning of many CM-2 and CM-5 machines, the best way to protect the substantial investment in CM Fortran codes is to port the codes to High Performance Fortran (HPF) on highly parallel systems. HPF is very similar to CM Fortran and thus represents a natural transition. Conversion issues involved in porting CM Fortran codes on the CM-5 to HPF are presented. In particular, the differences between data distribution directives and the CM Fortran Utility Routines Library, as well as the equivalent functionality in the HPF Library are discussed. Several CM Fortran codes (Cannon algorithm for matrix-matrix multiplication, Linear solver Ax=b, 1-D convolution for 2-D datasets, Laplace's Equation solver, and Direct Simulation Monte Carlo (DSMC) codes have been ported to Subset HPF on the IBM SP2 and the SGI Power Challenge. Speedup ratios versus number of processors for the Linear solver and DSMC code are presented.

Saini, Subhash↗