Engineering PapersSearch

NASA NTRS · 19960042903

Ensuring correct rollback recovery in distributed shared memory systems

Abstract

Distributed shared memory (DSM) implemented on a cluster of workstations is an increasingly attractive platform for executing parallel scientific applications. Checkpointing and rollback techniques can be used in such a system to allow the computation to progress in spite of the temporary failure of one or more processing nodes. This paper presents the design of an independent checkpointing method for DSM that takes advantage of DSM's specific properties to reduce error-free and rollback overhead. The scheme reduces the dependencies that need to be considered for correct rollback to those resulting from transfers of pages. Furthermore, in-transit messages can be recovered without the use of logging. We extend the scheme to a DSM implementation using lazy release consistency, where the frequency of dependencies is further reduced.

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Janssens, Bob, Fuchs, W. Kent. 1995-01-01. Ensuring correct rollback recovery in distributed shared memory systems. https://ntrs.nasa.gov/citations/19960042903

Cite the original work for its findings. Save a collection to share your selection of sources.