Multi-Bandit Best Arm Identification - Archive ouverte HAL Accéder directement au contenu
Rapport Année : 2011

Multi-Bandit Best Arm Identification

Victor Gabillon
  • Fonction : Auteur
  • PersonId : 900485
Mohammad Ghavamzadeh
  • Fonction : Auteur
  • PersonId : 868946
Alessandro Lazaric

Résumé

We study the problem of identifying the best arm in each of the bandits in a multi-bandit multi-armed setting. We first propose an algorithm called Gap-based Exploration (GapE) that focuses on the arms whose mean is close to the mean of the best arm in the same bandit (i.e., small gap). We then introduce an algorithm, called GapE-V, which takes into account the variance of the arms in addition to their gap. We prove an upper-bound on the probability of error for both algorithms. Since GapE and GapE-V need to tune an exploration parameter that depends on the complexity of the problem, which is often unknown in advance, we also introduce variation of these algorithms that estimates this complexity online. Finally, we evaluate the performance of these algorithms and compare them to other allocation strategies on a number of synthetic problems.

Domaines

Informatique
Fichier principal
Vignette du fichier
multi-bandit_techreport.pdf (279.09 Ko) Télécharger le fichier
Origine : Fichiers produits par l'(les) auteur(s)
Loading...

Dates et versions

hal-00632523 , version 1 (14-10-2011)
hal-00632523 , version 2 (25-10-2011)
hal-00632523 , version 3 (19-11-2011)

Identifiants

  • HAL Id : hal-00632523 , version 3

Citer

Victor Gabillon, Mohammad Ghavamzadeh, Alessandro Lazaric, Sébastien Bubeck. Multi-Bandit Best Arm Identification. 2011. ⟨hal-00632523v3⟩
297 Consultations
161 Téléchargements

Partager

Gmail Facebook X LinkedIn More