Active Optimistic Message Logging for Reliable Execution of MPI Applications - Archive ouverte HAL Accéder directement au contenu
Communication Dans Un Congrès Année : 2009

Active Optimistic Message Logging for Reliable Execution of MPI Applications

Résumé

To execute MPI applications reliably, fault tolerance mechanisms are needed. Message logging is a well known solution to provide fault tolerance for MPI applications. It as been proved that it can tolerate higher failure rate than coordinated checkpointing. However pessimistic and causal message logging can induce high overhead on failure free execution. In this paper, we present O2P, a new optimistic message logging protocol, based on active optimistic message logging. Contrary to existing optimistic message logging protocols that saves dependency information on reliable storage periodically, O2P logs dependency information as soon as possible to reduce the amount of data piggybacked on application messages. Thus it reduces the overhead of the protocol on failure free execution, making it more scalable and simplifying recovery. O2P is implemented as a module of the Open MPI library. Experiments show that active message logging is promising to improve scalability and performance of optimistic message logging.
Fichier non déposé

Dates et versions

inria-00424002 , version 1 (13-10-2009)

Identifiants

  • HAL Id : inria-00424002 , version 1

Citer

Thomas Ropars, Christine Morin. Active Optimistic Message Logging for Reliable Execution of MPI Applications. 15th International Euro-Par Conference, Aug 2009, Delft, Netherlands. ⟨inria-00424002⟩
230 Consultations
0 Téléchargements

Partager

Gmail Facebook X LinkedIn More