Skip to Main content Skip to Navigation
Conference papers

Multiple-Play Bandits in the Position-Based Model

Abstract : Sequentially learning to place items in multi-position displays or lists is a task that can be cast into the multiple-play semi-bandit setting. However, a major concern in this context is when the system cannot decide whether the user feedback for each item is actually exploitable. Indeed, much of the content may have been simply ignored by the user. The present work proposes to exploit available information regarding the display position bias under the so-called Position-based click model (PBM). We first discuss how this model differs from the Cascade model and its variants considered in several recent works on multiple-play bandits. We then provide a novel regret lower bound for this model as well as computationally efficient algorithms that display good empirical and theoretical performance.
Keywords : multi-armed bandits
Complete list of metadata

https://hal.archives-ouvertes.fr/hal-01328003
Contributor : Claire Vernade <>
Submitted on : Tuesday, June 7, 2016 - 1:24:24 PM
Last modification on : Tuesday, December 8, 2020 - 10:21:26 AM

Files

nips_hal.pdf
Files produced by the author(s)

Identifiers

  • HAL Id : hal-01328003, version 1
  • ARXIV : 1606.02448

Citation

Paul Lagrée, Claire Vernade, Olivier Cappé. Multiple-Play Bandits in the Position-Based Model. Neural Information Processing Systems (NIPS), Jan 2016, Barcelone, Spain. ⟨hal-01328003⟩

Share

Metrics

Record views

800

Files downloads

181