A Generic and High Performance Approach for Fault Tolerance in Communication Library - Archive ouverte HAL Access content directly
Reports (Research Report) Year : 2010

A Generic and High Performance Approach for Fault Tolerance in Communication Library

Abstract

With the increase of the number of nodes in clusters, the probability of failures increases. In this paper, we study the failures in the network stack for high performance networks. We present the design of several fault-tolerance mechanisms for communication libraries to detect failures and to ensure message integrity. We have implemented these mechanisms in the N EW M ADELEINE communication library with a quick detection of failures in a portable way, and with fallback to available links when an error occurs. Our mechanisms ensure the integrity of messages without lowering too much the networking performance. Our evaluation show that ensuring fault-tolerance does not impact significantly the performance of most applications.
Fichier principal
Vignette du fichier
main.pdf (245.29 Ko) Télécharger le fichier
Origin : Files produced by the author(s)
Loading...

Dates and versions

hal-00793176 , version 1 (22-02-2013)

Identifiers

  • HAL Id : hal-00793176 , version 1

Cite

François Trahay, Alexandre Denis, Yutaka Ishikawa. A Generic and High Performance Approach for Fault Tolerance in Communication Library. [Research Report] INRIA Bordeaux. 2010. ⟨hal-00793176⟩
145 View
132 Download

Share

Gmail Facebook X LinkedIn More