Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel machines”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Parallel solution of triangular systems of equations

The solution on parallel computers of systems of equations Lx = b, where L is lower triangular, is considered. Some authors have suggested that it is difficult to solve such systems in parallel on message-passing, local-memory machines when L is stored by columns. It is shown here that this is not necessarily the case if the machine can accomplish fan-in communication with reasonable efficiency.

Romine, Charles H.↗

Distributed intelligence for supervisory control

Supervisory control systems must deal with various types of intelligence distributed throughout the layers of control. Typical layers are real-time servo control, off-line planning and reasoning subsystems and finally, the human operator. Design methodologies must account for the fact that the majority of the intelligence will reside with the human operator. Hierarchical decompositions and feedback loops as conceptual building blocks that provide a common ground for man-machine interaction are discussed. Examples of types of parallelism and parallel implementation on several classes of computer architecture are also discussed.

Wolfe, W. J.↗

Distributed control architecture for real-time telerobotic operation

The emerging field of telerobotics places new demands on control system architecture to allow both autonomous operations and natural human-machine interfacing. The feasibility of multiprocessor systems performing parallel control computations is realizable. A practical distribution of control processors is presented and the issues involved in the realization of this architecture are discussed. A prototype dual axis controller based on the NOVIX computer is described, and results of its implementation are discussed. Application of this type of control system to a replicated, redundant manipulator system is also described.

Martin, H. L.↗

Motion detection in astronomical and ice floe images

Two approaches are presented for establishing correspondence between small areas in pairs of successive images for motion detection. The first one, based on local correlation, is used on a pair of successive Voyager images of the Jupiter which differ mainly in locally variable translations. This algorithm is implemented on a sequential machine (VAX 780) as well as the Massively Parallel Processor (MPP). In the case of the sequential algorithm, the pixel correspondence or match is computed on a sparse grid of points using nonoverlapping windows (typically 11 x 11) by local correlations over a predetermined search area. The displacement of the corresponding pixels in the two images is called the disparities to cubic surfaces. The disparities at points where the error between the computed values and the surface values exceeds a particular threshold are replaced by the surface values. A bilinear interpolation is then used to estimate disparities at all other pixels between the grid points. When this algorithm was applied at the red spot in the Jupiter image, the rotating velocity field of the storm was determined. The second method of motion detection is applicable to pairs of images in which corresponding areas can experience considerable translation as well as rotation.

Manohar, M.↗

The alignment-distribution graph

Implementing a data-parallel language such as Fortran 90 on a distributed-memory parallel computer requires distributing aggregate data objects (such as arrays) among the memory modules attached to the processors. The mapping of objects to the machine determines the amount of residual communication needed to bring operands of parallel operations into alignment with each other. We present a program representation called the alignment distribution graph that makes these communication requirements explicit. We describe the details of the representation, show how to model communication cost in this framework, and outline several algorithms for determining object mappings that approximately minimize residual communication.

Chatterjee, Siddhartha↗

The alignment-distribution graph

Implementing a data-parallel language such as Fortran 90 on a distributed-memory parallel computer requires distributing aggregate data objects (such as arrays) among the memory modules attached to the processors. The mapping of objects to the machine determines the amount of residual communication needed to bring operands of parallel operations into alignment with each other. We present a program representation called the alignment-distribution graph that makes these communication requirements explicit. We describe the details of the representation, show how to model communication cost in this framework, and outline several algorithms for determining object mappings that approximately minimize residual communication.

Chatterjee, Siddhartha↗

Performance Comparison of HPF and MPI Based NAS Parallel Benchmarks

Compilers supporting High Performance Form (HPF) features first appeared in late 1994 and early 1995 from Applied Parallel Research (APR), Digital Equipment Corporation, and The Portland Group (PGI). IBM introduced an HPF compiler for the IBM RS/6000 SP2 in April of 1996. Over the past two years, these implementations have shown steady improvement in terms of both features and performance. The performance of various hardware/ programming model (HPF and MPI) combinations will be compared, based on latest NAS Parallel Benchmark results, thus providing a cross-machine and cross-model comparison. Specifically, HPF based NPB results will be compared with MPI based NPB results to provide perspective on performance currently obtainable using HPF versus MPI or versus hand-tuned implementations such as those supplied by the hardware vendors. In addition, we would also present NPB, (Version 1.0) performance results for the following systems: DEC Alpha Server 8400 5/440, Fujitsu CAPP Series (VX, VPP300, and VPP700), HP/Convex Exemplar SPP2000, IBM RS/6000 SP P2SC node (120 MHz), NEC SX-4/32, SGI/CRAY T3E, and SGI Origin2000. We would also present sustained performance per dollar for Class B LU, SP and BT benchmarks.

Saini, Subhash↗

Issues on Reproducibility/Reliability of Magnetic NDE Methods

One of the critical elements related to the practicality of any NDE technique is its reproducibility under nominally the same inspection conditions. The results of certain test methodologies, however, are not always repeatable and understanding the origin of the irreproducibility is often as critical as obtaining reproducible results. One example is the characterization of residual stress in structural ferromagnets using the magnetoacoustic (MAC) method. Although it has not been widely publicized, the test results of this method are known to be time-dependent. Two distinct types of time dependencies have been observed during testing. The first type has a clearly definable relaxation time, while no such trend has been observed for the second. The purpose of the present study is to systematically investigate the time dependence of the second type, to find out the range and, if possible, the origin of the variation in the test results. For this, MAC curves were obtained under various stress levels and the tests were repeated over time. Particular attention was given to whether noise in the measuring device or a change in the laboratory environment could have been a contributing factor. The steel samples used for the study were cut from C- and U-class railroad wheels. The MAC behavior of these samples was reported previously. Each steel sample was first machined to be a cylindrical rod of 3.175 cm (1.25 in) in diameter and 26.67 cm (10.5 in) in length. The center portion of the samples were further machined to form a pair of flat and parallel surfaces for the launch and reflection of the ultrasonic pulses. Data acquisition involved two major elements; magnetic and acoustic measurements. Throughout the experiment the net magnetic induction, B, was measured by integrating the induction pickup coil output using an integrating fluxmeter. The acoustic measurements employed the phase-locked technique which will be described in the following.

Namkung, M.↗

Fluid dynamics parallel computer development at NASA Langley Research Center

To accomplish more detailed simulations of highly complex flows, such as the transition to turbulence, fluid dynamics research requires computers much more powerful than any available today. Only parallel processing on multiple-processor computers offers hope for achieving the required effective speeds. Looking ahead to the use of these machines, the fluid dynamicist faces three issues: algorithm development for near-term parallel computers, architecture development for future computer power increases, and assessment of possible advantages of special purpose designs. Two projects at NASA Langley address these issues. Software development and algorithm exploration is being done on the FLEX/32 Parallel Processing Research Computer. New architecture features are being explored in the special purpose hardware design of the Navier-Stokes Computer. These projects are complementary and are producing promising results.

Townsend, James C.↗

Mobile and replicated alignment of arrays in data-parallel programs

When a data-parallel language like FORTRAN 90 is compiled for a distributed-memory machine, aggregate data objects (such as arrays) are distributed across the processor memories. The mapping determines the amount of residual communication needed to bring operands of parallel operations into alignment with each other. A common approach is to break the mapping into two stages: first, an alignment that maps all the objects to an abstract template, and then a distribution that maps the template to the processors. We solve two facets of the problem of finding alignments that reduce residual communication: we determine alignments that vary in loops, and objects that should have replicated alignments. We show that loop-dependent mobile alignment is sometimes necessary for optimum performance, and we provide algorithms with which a compiler can determine good mobile alignments for objects within do loops. We also identify situations in which replicated alignment is either required by the program itself (via spread operations) or can be used to improve performance. We propose an algorithm based on network flow that determines which objects to replicate so as to minimize the total amount of broadcast communication in replication. This work on mobile and replicated alignment extends our earlier work on determining static alignment.

Chatterjee, Siddhartha↗

Methodologies and Tools for Tuning Parallel Programs: Facts and Fantasies

The need for computing power has forced a migration from serial computation on a single processor to parallel processing on multiprocessors. However, without effective means to monitor (and analyze) program execution, tuning the performance of parallel programs becomes exponentially difficult as program complexity and machine size increase. The recent introduction of performance tuning tools from various supercomputer vendors (Intel's ParAide, TMC's PRISM, CRI's Apprentice, and Convex's CXtrace) seems to indicate the maturity of performance tool technologies and vendors'/customers' recognition of their importance. However, a few important questions remain: What kind of performance bottlenecks can these tools detect (or correct)? How time consuming is the performance tuning process? What are some important technical issues that remain to be tackled in this area? This workshop reviews the fundamental concepts involved in analyzing and improving the performance of parallel and heterogeneous message-passing programs. Several alternative strategies will be contrasted, and for each we will describe how currently available tuning tools (e.g. AIMS, ParAide, PRISM, Apprentice, CXtrace, ATExpert, Pablo, IPS-2) can be used to facilitate the process. We will characterize the effectiveness of the tools and methodologies based on actual user experiences at NASA Ames Research Center. Finally, we will discuss their limitations and outline recent approaches taken by vendors and the research community to address them.

Yan, Jerry C.↗

Performance Evaluation Methodologies and Tools for Massively Parallel Programs

The need for computing power has forced a migration from serial computation on a single processor to parallel processing on multiprocessors. However, without effective means to monitor (and analyze) program execution, tuning the performance of parallel programs becomes exponentially difficult as program complexity and machine size increase. The recent introduction of performance tuning tools from various supercomputer vendors (Intel's ParAide, TMC's PRISM, CSI'S Apprentice, and Convex's CXtrace) seems to indicate the maturity of performance tool technologies and vendors'/customers' recognition of their importance. However, a few important questions remain: What kind of performance bottlenecks can these tools detect (or correct)? How time consuming is the performance tuning process? What are some important technical issues that remain to be tackled in this area? This workshop reviews the fundamental concepts involved in analyzing and improving the performance of parallel and heterogeneous message-passing programs. Several alternative strategies will be contrasted, and for each we will describe how currently available tuning tools (e.g., AIMS, ParAide, PRISM, Apprentice, CXtrace, ATExpert, Pablo, IPS-2)) can be used to facilitate the process. We will characterize the effectiveness of the tools and methodologies based on actual user experiences at NASA Ames Research Center. Finally, we will discuss their limitations and outline recent approaches taken by vendors and the research community to address them.

Yan, Jerry C.↗

Methodologies and Tools for Tuning Parallel Programs: 80% Art, 20% Science, and 10% Luck

The need for computing power has forced a migration from serial computation on a single processor to parallel processing on multiprocessors. However, without effective means to monitor (and analyze) program execution, tuning the performance of parallel programs becomes exponentially difficult as program complexity and machine size increase. In the past few years, the ubiquitous introduction of performance tuning tools from various supercomputer vendors (Intel's ParAide, TMC's PRISM, CRI's Apprentice, and Convex's CXtrace) seems to indicate the maturity of performance instrumentation/monitor/tuning technologies and vendors'/customers' recognition of their importance. However, a few important questions remain: What kind of performance bottlenecks can these tools detect (or correct)? How time consuming is the performance tuning process? What are some important technical issues that remain to be tackled in this area? This workshop reviews the fundamental concepts involved in analyzing and improving the performance of parallel and heterogeneous message-passing programs. Several alternative strategies will be contrasted, and for each we will describe how currently available tuning tools (e.g. AIMS, ParAide, PRISM, Apprentice, CXtrace, ATExpert, Pablo, IPS-2) can be used to facilitate the process. We will characterize the effectiveness of the tools and methodologies based on actual user experiences at NASA Ames Research Center. Finally, we will discuss their limitations and outline recent approaches taken by vendors and the research community to address them.

Yan, Jerry C.↗

Diamond Machining of an Off-Axis Biconic Aspherical Mirror

Two diamond-machining methods have been developed as part of an effort to design and fabricate an off-axis, biconic ellipsoidal, concave aluminum mirror for an infrared spectrometer at the Kitt Peak National Observatory. Beyond this initial application, the methods can be expected to enable satisfaction of requirements for future instrument mirrors having increasingly complex (including asymmetrical), precise shapes that, heretofore, could not readily be fabricated by diamond machining or, in some cases, could not be fabricated at all. In the initial application, the mirror is prescribed, in terms of Cartesian coordinates x and y, by aperture dimensions of 94 by 76 mm, placements of -2 mm off axis in x and 227 mm off axis in y, an x radius of curvature of 377 mm, a y radius of curvature of 407 mm, an x conic constant of 0.078, and a y conic constant of 0.127. The aspect ratio of the mirror blank is about 6. One common, "diamond machining" process uses single-point diamond turning (SPDT). However, it is impossible to generate the required off-axis, biconic ellipsoidal shape by conventional SPDT because (1) rotational symmetry is an essential element of conventional SPDT and (2) the present off-axis biconic mirror shape lacks rotational symmetry. Following conventional practice, it would be necessary to make this mirror from a glass blank by computer-controlled polishing, which costs more than diamond machining and yields a mirror that is more difficult to mount to a metal bench. One of the two present diamond machining methods involves the use of an SPDT machine equipped with a fast tool servo (FTS). The SPDT machine is programmed to follow the rotationally symmetric asphere that best fits the desired off-axis, biconic ellipsoidal surface. The FTS is actuated in synchronism with the rotation of the SPDT machine to generate the difference between the desired surface and the best-fit rotationally symmetric asphere. In order to minimize the required stroke of the FTS, the blanks were positioned at a large off-axis distance and angle, and the axis of the FTS was not parallel to the axis of the spindle of the SPDT machine. The spindle was rotated at a speed of 120 rpm, and the maximum FTS speed was 8.2 mm/s.

Ohl, Raymond G.↗

Large space structures fabrication experiment

The fabrication machine used for the rolltrusion and on-orbit forming of graphite thermoplastic (CTP) strip material into structural sections is described. The basic process was analytically developed parallel with, and integrated into the conceptual design of, a flight experiment machine for producing a continuous triangular cross section truss. The machine and its associated ancillary equipment are mounted on a Space Lab pallet. Power, thermal control, and instrumentation connections are made during ground installation. Observation, monitoring, caution and warning, and control panels and displays are installed at the payload specialist station in the orbiter. The machine is primed before flight by initiation of beam forming, to include attachment of the first set of cross members and anchoring of the diagonal cords. Control of the experiment will be from the orbiter mission specialist station. Normal operation is by automatic processing control software. Machine operating data are displayed and recorded on the ground. Data is processed and formatted to show progress of the major experiment parameters including stable operation, physical symmetry, joint integrity, and structural properties.

Source record↗

Performance Evaluation of Three Distributed Computing Environments for Scientific Applications

We present performance results for three distributed computing environments using the three simulated CFD applications in the NAS Parallel Benchmark suite. These environments are the DCF cluster, the LACE cluster, and an Intel iPSC/860 machine. The DCF is a prototypic cluster of loosely coupled SGI R3000 machines connected by Ethernet. The LACE cluster is a tightly coupled cluster of 32 IBM RS6000/560 machines connected by Ethernet as well as by either FDDI or an IBM Allnode switch. Results of several parallel algorithms for the three simulated applications are presented and analyzed based on the interplay between the communication requirements of an algorithm and the characteristics of the communication network of a distributed system.

Fatoohi, Rod↗

Automated solar module assembly line

The solar module assembly machine which Kulicke and Soffa delivered under this contract is a cell tabbing and stringing machine, and capable of handling a variety of cells and assembling strings up to 4 feet long which then can be placed into a module array up to 2 feet by 4 feet in a series of parallel arrangement, and in a straight or interdigitated array format. The machine cycle is 5 seconds per solar cell. This machine is primarily adapted to 3 inch diameter round cells with two tabs between cells. Pulsed heat is used as the bond technique for solar cell interconnects. The solar module assembly machine unloads solar cells from a cassette, automatically orients them, applies flux and solders interconnect ribbons onto the cells. It then inverts the tabbed cells, connects them into cell strings, and delivers them into a module array format using a track mounted vacuum lance, from which they are taken to test and cleaning benches prior to final encapsulation into finished solar modules. Throughout the machine the solar cell is handled very carefully, and any contact with the collector side of the cell is avoided or minimized.

Bycer, M.↗

Implementation and analysis of a Navier-Stokes algorithm on parallel computers

The results of the implementation of a Navier-Stokes algorithm on three parallel/vector computers are presented. The object of this research is to determine how well, or poorly, a single numerical algorithm would map onto three different architectures. The algorithm is a compact difference scheme for the solution of the incompressible, two-dimensional, time-dependent Navier-Stokes equations. The computers were chosen so as to encompass a variety of architectures. They are the following: the MPP, an SIMD machine with 16K bit serial processors; Flex/32, an MIMD machine with 20 processors; and Cray/2. The implementation of the algorithm is discussed in relation to these architectures and measures of the performance on each machine are given. The basic comparison is among SIMD instruction parallelism on the MPP, MIMD process parallelism on the Flex/32, and vectorization of a serial code on the Cray/2. Simple performance models are used to describe the performance. These models highlight the bottlenecks and limiting factors for this algorithm on these architectures. Finally, conclusions are presented.

Fatoohi, Raad A.↗