Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “IMPLEMENTATION”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Integrated System Health Management: Foundational Concepts, Approach, and Implementation.

Implementation of integrated system health management (ISHM) capability is fundamentally linked to the management of data, information, and knowledge (DIaK) with the purposeful objective of determining the health of a system. It is akin to having a team of experts who are all individually and collectively observing and analyzing a complex system, and communicating effectively with each other in order to arrive to an accurate and reliable assessment of its health. We present concepts, procedures, and a specific approach as a foundation for implementing a credible ISHM capability. The capability stresses integration of DIaK from all elements of a system. The intent is also to make possible implementation of on-board ISHM capability, in contrast to a remote capability. The information presented is the result of many years of research, development, and maturation of technologies, and of prototype implementations in operational systems (rocket engine test facilities). The paper will address the following topics: 1. ISHM Model of a system 2. Detection of anomaly indicators. 3. Determination and confirmation of anomalies. 4. Diagnostic of causes and determination of effects. 5. Consistency checking cycle. 6. Management of health information 7. User Interfaces 8. Example implementation ISHM has been defined from many perspectives. We define it as a capability that might be achieved by various approaches. We describe a specific approach that has been matured throughout many years of development, and pilot implementations. ISHM is a capability that is achieved by integrating data, information, and knowledge (DIaK) that might be distributed throughout the system elements (which inherently implies capability to manage DIaK associated with distributed sub-systems). DIaK must be available to any element of a system at the right time and in accordance with a meaningful context. ISHM Functional Capability Level (FCL) is measured by how well a system performs the following functions: (1) detect anomalies, (2) diagnose causes, (3) predict future anomalies/failures, and (4) provide the user with an integrated awareness about the condition of every element in the system and guide user decisions.

Figueroa, Fernando↗

Hardware Implementation of a Bilateral Subtraction Filter

A bilateral subtraction filter has been implemented as a hardware module in the form of a field-programmable gate array (FPGA). In general, a bilateral subtraction filter is a key subsystem of a high-quality stereoscopic machine vision system that utilizes images that are large and/or dense. Bilateral subtraction filters have been implemented in software on general-purpose computers, but the processing speeds attainable in this way even on computers containing the fastest processors are insufficient for real-time applications. The present FPGA bilateral subtraction filter is intended to accelerate processing to real-time speed and to be a prototype of a link in a stereoscopic-machine- vision processing chain, now under development, that would process large and/or dense images in real time and would be implemented in an FPGA. In terms that are necessarily oversimplified for the sake of brevity, a bilateral subtraction filter is a smoothing, edge-preserving filter for suppressing low-frequency noise. The filter operation amounts to replacing the value for each pixel with a weighted average of the values of that pixel and the neighboring pixels in a predefined neighborhood or window (e.g., a 9 9 window). The filter weights depend partly on pixel values and partly on the window size. The present FPGA implementation of a bilateral subtraction filter utilizes a 9 9 window. This implementation was designed to take advantage of the ability to do many of the component computations in parallel pipelines to enable processing of image data at the rate at which they are generated. The filter can be considered to be divided into the following parts (see figure): a) An image pixel pipeline with a 9 9- pixel window generator, b) An array of processing elements; c) An adder tree; d) A smoothing-and-delaying unit; and e) A subtraction unit. After each 9 9 window is created, the affected pixel data are fed to the processing elements. Each processing element is fed the pixel value for its position in the window as well as the pixel value for the central pixel of the window. The absolute difference between these two pixel values is calculated and used as an address in a lookup table. Each processing element has a lookup table, unique for its position in the window, containing the weight coefficients for the Gaussian function for that position. The pixel value is multiplied by the weight, and the outputs of the processing element are the weight and pixel-value weight product. The products and weights are fed to the adder tree. The sum of the products and the sum of the weights are fed to the divider, which computes the sum of products the sum of weights. The output of the divider is denoted the bilateral smoothed image. The smoothing function is a simple weighted average computed over a 3 3 subwindow centered in the 9 9 window. After smoothing, the image is delayed by an additional amount of time needed to match the processing time for computing the bilateral smoothed image. The bilateral smoothed image is then subtracted from the 3 3 smoothed image to produce the final output. The prototype filter as implemented in a commercially available FPGA processes one pixel per clock cycle. Operation at a clock speed of 66 MHz has been demonstrated, and results of a static timing analysis have been interpreted as suggesting that the clock speed could be increased to as much as 100 MHz.

Huertas, Andres↗

Optimized FPGA Implementation of Multi-Rate FIR Filters Through Thread Decomposition

Multi-rate finite impulse response (MRFIR) filters are among the essential signal-processing components in spaceborne instruments where finite impulse response filters are often used to minimize nonlinear group delay and finite precision effects. Cascaded (multistage) designs of MRFIR filters are further used for large rate change ratio in order to lower the required throughput, while simultaneously achieving comparable or better performance than single-stage designs. Traditional representation and implementation of MRFIR employ polyphase decomposition of the original filter structure, whose main purpose is to compute only the needed output at the lowest possible sampling rate. In this innovation, an alternative representation and implementation technique called TD-MRFIR (Thread Decomposition MRFIR) is presented. The basic idea is to decompose MRFIR into output computational threads, in contrast to a structural decomposition of the original filter as done in the polyphase decomposition. A naive implementation of a decimation filter consisting of a full FIR followed by a downsampling stage is very inefficient, as most of the computations performed by the FIR state are discarded through downsampling. In fact, only 1/M of the total computations are useful (M being the decimation factor). Polyphase decomposition provides an alternative view of decimation filters, where the downsampling occurs before the FIR stage, and the outputs are viewed as the sum of M sub-filters with length of N/M taps. Although this approach leads to more efficient filter designs, in general the implementation is not straightforward if the numbers of multipliers need to be minimized. In TD-MRFIR, each thread represents an instance of the finite convolution required to produce a single output of the MRFIR. The filter is thus viewed as a finite collection of concurrent threads. Each of the threads completes when a convolution result (filter output value) is computed, and activated when the first input of the convolution becomes available. Thus, the new threads get spawned at exactly the rate of N/M, where N is the total number of taps, and M is the decimation factor. Existing threads retire at the same rate of N/M. The implementation of an MRFIR is thus transformed into a problem to statically schedule the minimum number of multipliers such that all threads can be completed on time. Solving the static scheduling problem is rather straightforward if one examines the Thread Decomposition Diagram, which is a table-like diagram that has rows representing computation threads and columns representing time. The control logic of the MRFIR can be implemented using simple counters. Instead of decomposing MRFIRs into subfilters as suggested by polyphase decomposition, the thread decomposition diagrams transform the problem into a familiar one of static scheduling, which can be easily solved as the input rate is constant.

Kobayashi, Kayla N.↗

Performance Analysis of a Hardware Implemented Complex Signal Kurtosis Radio-Frequency Interference Detector

Radio-frequency interference (RFI) is a known problem for passive remote sensing as evidenced in the L-band radiometers SMOS, Aquarius and more recently, SMAP. Various algorithms have been developed and implemented on SMAP to improve science measurements. This was achieved by the use of a digital microwave radiometer. RFI mitigation becomes more challenging for microwave radiometers operating at higher frequencies in shared allocations. At higher frequencies larger bandwidths are also desirable for lower measurement noise further adding to processing challenges. This work focuses on finding improved RFI mitigation techniques that will be effective at additional frequencies and at higher bandwidths. To aid the development and testing of applicable detection and mitigation techniques, a wide-band RFI algorithm testing environment has been developed using the Reconfigurable Open Architecture Computing Hardware System (ROACH) built by the Collaboration for Astronomy Signal Processing and Electronics Research (CASPER) Group. The testing environment also consists of various test equipment used to reproduce typical signals that a radiometer may see including those with and without RFI. The testing environment permits quick evaluations of RFI mitigation algorithms as well as show that they are implementable in hardware. The algorithm implemented is a complex signal kurtosis detector which was modeled and simulated. The complex signal kurtosis detector showed improved performance over the real kurtosis detector under certain conditions. The real kurtosis is implemented on SMAP at 24 MHz bandwidth. The complex signal kurtosis algorithm was then implemented in hardware at 200 MHz bandwidth using the ROACH. In this work, performance of the complex signal kurtosis and the real signal kurtosis are compared. Performance evaluations and comparisons in both simulation as well as experimental hardware implementations were done with the use of receiver operating characteristic (ROC) curves. The complex kurtosis algorithm has the potential to reduce data rate due to onboard processing in addition to improving RFI detection performance.

microwaves↗

Performance Analysis of a Hardware Implemented Complex Signal Kurtosis Radio-Frequency Interference Detector

Radio-frequency interference (RFI) is a known problem for passive remote sensing as evidenced in the L-band radiometers SMOS, Aquarius and more recently, SMAP. Various algorithms have been developed and implemented on SMAP to improve science measurements. This was achieved by the use of a digital microwave radiometer. RFI mitigation becomes more challenging for microwave radiometers operating at higher frequencies in shared allocations. At higher frequencies larger bandwidths are also desirable for lower measurement noise further adding to processing challenges. This work focuses on finding improved RFI mitigation techniques that will be effective at additional frequencies and at higher bandwidths. To aid the development and testing of applicable detection and mitigation techniques, a wide-band RFI algorithm testing environment has been developed using the Reconfigurable Open Architecture Computing Hardware System (ROACH) built by the Collaboration for Astronomy Signal Processing and Electronics Research (CASPER) Group. The testing environment also consists of various test equipment used to reproduce typical signals that a radiometer may see including those with and without RFI. The testing environment permits quick evaluations of RFI mitigation algorithms as well as show that they are implementable in hardware. The algorithm implemented is a complex signal kurtosis detector which was modeled and simulated. The complex signal kurtosis detector showed improved performance over the real kurtosis detector under certain conditions. The real kurtosis is implemented on SMAP at 24 MHz bandwidth. The complex signal kurtosis algorithm was then implemented in hardware at 200 MHz bandwidth using the ROACH. In this work, performance of the complex signal kurtosis and the real signal kurtosis are compared. Performance evaluations and comparisons in both simulation as well as experimental hardware implementations were done with the use of receiver operating characteristic (ROC) curves.

radiometers↗

Implementation of Topology Optimization Using OpenMDAO

Topology optimization is one of the widely known branches among the structural optimization, and it distinguishes itself by a large design domain and versatility. It can determine the optimal design out of an infinite number of configurations, thereby drawing interest from both industry and academia with regards to its applicability to additive manufacturing. However, its implementation is often a daunting task for engineers in practice. One of the issues is the programming effort required whenever the implementation requires any changes, ranging from subtle tweaks to drastic changes in the problem definition, and the derivative computation must be correspondingly updated after any change. The implementation, therefore, is not only time-consuming but also repetitive and susceptible to human-induced errors. In this regard, topology optimization implementations stand to benefit from reusability, ease of restructuring, and modularity. In this work, we propose OpenMDAO, a computational framework for multidisciplinary design optimization (MDO), as a generic platform for topology optimization. Two widely used topology optimization techniquesâ€"density-based and level-setâ€"are implemented as a demonstration. These techniques are implemented in a decomposed manner, with the aid of the modular architecture of OpenMDAO as well as state-of-the-art numerical methods. Variations on the density-based topology optimization approach are shown to demonstrate the modularity and automation for derivative computation that OpenMDAO provides.

Chung, Hayoung↗

Secondary Science Teachers’ Implementation of a Curricular Intervention When Teaching With Global Climate Models

In the past decade, emphasis on promoting “climate literacy” in K-16 science classrooms has increased. Teachers play a critical role in cultivating these opportunities, especially in secondary science classrooms. However, most prior climate education research has focused on students and student learning; little is known about how teachers implement climate-focused curricular interventions. Here, we report findings from a concurrent mixed methods, multiple-case study of four secondary science teachers’ implementation of a new, NGSS-aligned, model-centric climate curriculum module grounded in the use of a data-driven, computer-based climate modeling tool—Easy Global Climate Model (EzGCM). We employ multiple data sources, including video-recorded classroom observations, interviews, and instructional artifacts, and both qualitative and quantitative analyses, to investigate how teachers implemented the curriculum. Findings show that, overall, teachers implemented the curriculum in ways that were less model-centric than designed, placing greater emphasis on EzGCM itself rather than using the model to investigate Earth’s changing climate. Additionally, we present detailed single-case studies of each participant teacher that highlight differences in teachers’ implementation of the curriculum module and their reasoning for making observed instructional decisions. This research sheds light on the design of secondary science learning environments by illustrating the varied ways teachers implement a climate-focused curriculum to support students’ developing climate literacy. This has important implications for the design of climate-focused curriculum and supports for teachers.

Secondary science teaching↗

BSE versus StarTrack: Implementations of new wind, remnant-formation, and natal-kick schemes in NBODY7 and their astrophysical consequences

As a result of their formation via massive single and binary stellar evolution, the masses of stellar-remnant black holes (BH) are subjects of great interest in this era of gravitational-wave detection from binary black hole (BBH) and binary neutron star merger events. In this work, we present new developments in the stellar-remnant formation and related schemes of the current N-body evolution program NBODY7. We demonstrate that the newly implemented stellar-wind and remnant-formation schemes in the stellar-evolutionary sector or BSE of the NBODY7 code, such as the “rapid” and the “delayed” supernova (SN) schemes along with an implementation of pulsational-pair-instability and pair-instability supernova (PPSN/PSN), now produce neutron star (NS) and BH masses that agree nearly perfectly, over large ranges of zero-age-main-sequence (ZAMS) mass and metallicity, with those from the widely recognised StarTrack population-synthesis program. We also demonstrate the new, recipe-based implementations of various widely debated mechanisms of natal kicks on NSs and BHs, such as “convection-asymmetry-driven”, “collapse-asymmetry-driven”, and “neutrino-emission-driven” kicks, in addition to a fully consistent implementation of the standard, fallback-dependent, momentum-conserving natal kick. All the above newly implemented schemes are also shared with the standalone versions of SSE and BSE. Furthermore, all these demonstrations are performed with both the updated standalone BSE and the updated NBODY7/BSE. When convolved with stellar and primordial-binary populations as observed in young massive clusters, such remnant-formation and natal-kick mechanisms crucially determine the accumulated number, mass, and mass distribution of the BHs retained in young massive, open, and globular clusters (GCs); these BHs would eventually become available for long-term dynamical processing. Among other conclusions, we find that although the newer, delayed SN remnant formation model gives birth to the largest number (mass) of BHs, the older remnant-formation schemes cause the largest number (mass) of BHs to survive in clusters, when incorporating SN material fallback onto the BHs. The SN material fallback also causes the convection-asymmetry-driven SN kick to effectively retain similar numbers and masses of BHs in clusters as for the standard, momentum-conserving kick. The collapse-asymmetry-driven SN kick would cause nearly all BHs to be retained in clusters irrespective of their mass, remnant-formation model, and metallicity, whereas the inference of a large population of BHs in GCs would potentially rule out the neutrino-driven SN kick mechanism. Pre-SN mergers of massive primordial binaries would potentially cause BH masses to deviate from the theoretical, single-star ZAMS to mass-remnant mass relation unless a substantial of the total merging stellar mass of up to ≈40% is lost during a merger process. In particular, such mergers, at low metallicities, have the potential to produce low-spinning BHs within the PSN mass gap that can be retained in a stellar cluster and be available for subsequent dynamical interactions. As recent studies indicate, the new remnant-formation modelling reassures us that young massive and open clusters would potentially contribute to the dynamical BBH merger detection rate to a similar extent as their more massive GC counterparts.

79 ASTRONOMY AND ASTROPHYSICS↗

A Simple, Scalable Large Deformation Solid Mechanics Implementation in the MOOSE Framework

This article describes a large deformation solid mechanics solver implemented as part of the freely available and open source MOOSE finite element simulation framework. The article documents the choices made in developing the solid mechanics framework and describes novel formulations for the gradient operator and constitutive modeling framework made to simplify implementations of different coordinate systems, stabilized gradient operators, and different constitutive model inputs and outputs. In the process, the article describes a new formulation that casts objective integration of the Cauchy stress as a linear transformation of the small stress rate. Finally, the article presents key implementation details and examines the parallel efficiency of the solid mechanics solver implemented in MOOSE. The implementation retains a good weak scaling efficiency beyond 1,000 parallel processes. The article includes a discussion of the factors limiting the parallel efficiency of implicit, large deformation solid mechanics codes on current high-performance computers, with the main current limitation being the scalability of the algebraic multigrid methods used to solve the linearized equilibrium equations.

Applied computing → Computer-aided design↗

Implementation of Perturbation Theory and Sensitivity Capabilities in Griffin

Griffin is a Multiphysics Object-Oriented Simulation Environment (MOOSE) based reactor Multiphysics analysis application, jointly developed by Argonne and Idaho National Laboratories under the DOE-NE NEAMS program. This fiscal year, capabilities for reactivity and sensitivity evaluation using perturbation methods were implemented and verified. The First Order Perturbation Method (FOPT) was employed to compute reactivity worth resulting from small perturbations in input parameters, while the Generalized Perturbation Theory (GPT) was used to evaluate sensitivities of a range of response types, including reaction rate ratio, k-eigenvalue, neutron generation time, and effective delayed neutron fraction. These perturbation methods enable users to quantify how response quantities change due to a perturbation in a input parameter without explicitly performing an additional transport simulation for each perturbed state. In particular, the GPT formulation accounts for indirect effects arising from flux changes by solving generalized inhomogeneous equations, for which a Neumann series-based iterative solution method was developed and implemented in Griffin. The implemented reactivity and sensitivity evaluation capabilities were verified using two test problems: an infinite homogeneous system and a two-dimensional hexagonal core. The results showed excellent agreement with reference solutions obtained by a direct method based on finite difference approximation as well as GPT-based results from the PERSENT code, confirming the accuracy of both reactivity and sensitivity evaluations. Additionally, preliminary uncertainty quantification (UQ) results were obtained by combining the sensitivity values computed using GPT and external covariance data, demonstrating that the implemented sensitivity results can be reliably used for uncertainty calculations. To further demonstrate the generality and practical strength of the implementation, the sensitivity evaluation capability was successfully applied to the Empire microreactor with a geometrically complex design that poses significant modeling challenges. The results confirm that Griffin enables sensitivity evaluations even for irregular and highly heterogeneous reactor configurations, thereby establishing a foundation for UQ applications in advanced reactor designs and analyses.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Formal semantic specifications as implementation blueprints for real-time programming languages

Formal definitions of language and system semantics provide highly desirable checks on the correctness of implementations of programming languages and their runtime support systems. If these definitions can give concrete guidance to the implementor, major increases in implementation accuracy and decreases in implementation effort can be achieved. It is shown that of the wide variety of available methods the Hgraph (hypergraph) definitional technique (Pratt, 1975), is best suited to serve as such an implementation blueprint. A discussion and example of the Hgraph technique is presented, as well as an overview of the growing body of implementation experience of real-time languages based on Hgraph semantic definitions.

Feyock, S.↗

A real time microcomputer implementation of sensor failure detection for turbofan engines

An algorithm was developed which detects, isolates, and accommodates sensor failures using analytical redundancy. The performance of this algorithm was demonstrated on a full-scale F100 turbofan engine. The algorithm was implemented in real-time on a microprocessor-based controls computer which includes parallel processing and high order language programming. Parallel processing was used to achieve the required computational power for the real-time implementation. High order language programming was used in order to reduce the programming and maintenance costs of the algorithm implementation software. The sensor failure algorithm was combined with an existing multivariable control algorithm to give a complete control implementation with sensor analytical redundancy. The real-time microprocessor implementation of the algorithm which resulted in the successful completion of the algorithm engine demonstration, is described.

Delaat, John C.↗

A real time microcomputer implementation of sensor failure detection for turbofan engines

An algorithm was developed that detects, isolates, and accommodates sensor failures using analytical redundancy. The performance of this algorithm has been demonstrated on a full-scale F100 turbofan engine. The algorithm was implemented in real-time on a microprocessor-based controls computer which includes parallel processing and high order language programming. Parallel processing was used to achieve the required computational power for the real-time implementation. High order language programming was used in order to reduce the programming and maintenance costs of the algorithm implementation software. The sensor failure algorithm was combined with an existing multivariable control algorithm to give a complete control implementation with sensor analytical redundancy. The real-time microprocessor implementation of the algorithm, which resulted in the successful completion of the algorithm engine demonstration, is described.

Delaat, John C.↗

The effects of long delay and transmission errors on the performance of TP-4 implementations

A set of tools that allows us to measure and examine the effects of transmission delay and errors on the performance of TP-4 implementations has been developed. The tools give insight into both the large- and small-scale behaviors of an implementation. These tools have been systematically applied to a commercial implementation of TP-4. Measurements show, among other things, that a 2-second one-way transmission delay and an effective bit-error rate of 1 error per 100,000 bits can result in a 95 percent reduction in TP-4 throughput. The detailed statistics give insight into why transmission delay and errors affect this implementations so significantly and support a number of 'lessons learned' that could be applied to TP-4 implementations that operate more robustly across networks with long transmission delays and transmission errors.

Durst, Robert C.↗

Implementation of ADI: Schemes on MIMD parallel computers

In order to simulate the effects of the impingement of hot exhaust jets of High Performance Aircraft on landing surfaces a multi-disciplinary computation coupling flow dynamics to heat conduction in the runway needs to be carried out. Such simulations, which are essentially unsteady, require very large computational power in order to be completed within a reasonable time frame of the order of an hour. Such power can be furnished by the latest generation of massively parallel computers. These remove the bottleneck of ever more congested data paths to one or a few highly specialized central processing units (CPU's) by having many off-the-shelf CPU's work independently on their own data, and exchange information only when needed. During the past year the first phase of this project was completed, in which the optimal strategy for mapping an ADI-algorithm for the three dimensional unsteady heat equation to a MIMD parallel computer was identified. This was done by implementing and comparing three different domain decomposition techniques that define the tasks for the CPU's in the parallel machine. These implementations were done for a Cartesian grid and Dirichlet boundary conditions. The most promising technique was then used to implement the heat equation solver on a general curvilinear grid with a suite of nontrivial boundary conditions. Finally, this technique was also used to implement the Scalar Penta-diagonal (SP) benchmark, which was taken from the NAS Parallel Benchmarks report. All implementations were done in the programming language C on the Intel iPSC/860 computer.

Vanderwijngaart, Rob F.↗

Study of the Alsys implementation of the Catalogue of Interface Features and Options for the Ada language for 80386 Unix

A set of programs was written to test the functionality and performance of the Alsys Ada implementation of the Catalogue of Interface Features and Options (CIFO), a set of optional Ada packages for real-time applications. No problems were found with the task id, preemption control, or shared-data packages. Minor problems were found with the dispatching control, dynamic priority, events, non-waiting entry call, semaphore, and scheduling packages. The Alsys implementation is derived mostly from Release 2 of the CIFO standard, but includes some of the features of Release 3 and some modifications unique to Alsys. Performance measurements show that the semaphore and shared-data features are an order-of-magnitude faster than the same mechanisms using an Ada rendezvous. The non-waiting entry call is slightly faster than a standard rendezvous. The existence of errors in the implementation, the incompleteness of the documentation from the published standard impair the usefulness of this implementation. Despite those short-comings, the Alsys CIFO implementation might be of value in the development of real-time applications.

Gibson, James S.↗

Design, Implementation, and Verification of the Reliable Multicast Protocol

This document describes the Reliable Multicast Protocol (RMP) design, first implementation, and formal verification. RMP provides a totally ordered, reliable, atomic multicast service on top of an unreliable multicast datagram service. RMP is fully and symmetrically distributed so that no site bears an undue portion of the communications load. RMP provides a wide range of guarantees, from unreliable delivery to totally ordered delivery, to K-resilient, majority resilient, and totally resilient atomic delivery. These guarantees are selectable on a per message basis. RMP provides many communication options, including virtual synchrony, a publisher/subscriber model of message delivery, a client/server model of delivery, mutually exclusive handlers for messages, and mutually exclusive locks. It has been commonly believed that total ordering of messages can only be achieved at great performance expense. RMP discounts this. The first implementation of RMP has been shown to provide high throughput performance on Local Area Networks (LAN). For two or more destinations a single LAN, RMP provides higher throughput than any other protocol that does not use multicast or broadcast technology. The design, implementation, and verification activities of RMP have occurred concurrently. This has allowed the verification to maintain a high fidelity between design model, implementation model, and the verification model. The restrictions of implementation have influenced the design earlier than in normal sequential approaches. The protocol as a whole has matured smoother by the inclusion of several different perspectives into the product development.

Montgomery, Todd L.↗

Performance Comparison of a Matrix Solver on a Heterogeneous Network Using Two Implementations of MPI: MPICH and LAM

Two of the current and most popular implementations of the Message-Passing Standard, Message Passing Interface (MPI), were contrasted: MPICH by Argonne National Laboratory, and LAM by the Ohio Supercomputer Center at Ohio State University. A parallel skyline matrix solver was adapted to be run in a heterogeneous environment using MPI. The Message-Passing Interface Forum was held in May 1994 which lead to a specification of library functions that implement the message-passing model of parallel communication. LAM, which creates it's own environment, is more robust in a highly heterogeneous network. MPICH uses the environment native to the machine architecture. While neither of these free-ware implementations provides the performance of native message-passing or vendor's implementations, MPICH begins to approach that performance on the SP-2. The machines used in this study were: IBM RS6000, 3 Sun4, SGI, and the IBM SP-2. Each machine is unique and a few machines required specific modifications during the installation. When installed correctly, both implementations worked well with only minor problems.

Phillips, Jennifer K.↗