Bridging the Gap Between Imitation Learning and Inverse Reinforcement Learning

Bilal Piot; Matthieu Geist; Olivier Pietquin

doi:10.1109/TNNLS.2016.2543000

Article Dans Une Revue IEEE Transactions on Neural Networks and Learning Systems Année : 2017

Bridging the Gap Between Imitation Learning and Inverse Reinforcement Learning

(1, 2) , (3, 4) , (2, 1)

1
2
3
4

Bilal Piot

Fonction : Auteur

DeepMind [London]

Sequential Learning

Matthieu Geist

Fonction : Auteur
PersonId : 6945
IdHAL : matthieu-geist

CentraleSupélec

Laboratoire Interdisciplinaire des Environnements Continentaux

Olivier Pietquin

Fonction : Auteur
PersonId : 4024
IdHAL : olivier-pietquin
ORCID : 0000-0002-5386-465X
IdRef : 142821861

Sequential Learning

DeepMind [London]

Résumé

—Learning from Demonstrations (LfD) is a paradigm by which an apprentice agent learns a control policy for a dynamic environment by observing demonstrations delivered by an expert agent. It is usually implemented as either Imitation Learning (IL) or Inverse Reinforcement Learning (IRL) in the literature. On the one hand, IRL is a paradigm relying on Markov Decision Processes (MDPs), where the goal of the apprentice agent is to find a reward function from the expert demonstrations that could explain the expert behavior. On the other hand, IL consists in directly generalizing the expert strategy, observed in the demonstrations, to unvisited states (and it is therefore close to classification, when there is a finite set of possible decisions). While these two visions are often considered as opposite to each other, the purpose of this paper is to exhibit a formal link between these approaches from which new algorithms can be derived. We show that IL and IRL can be redefined in a way that they are equivalent, in the sense that there exists an explicit bijective operator (namely the inverse optimal Bellman operator) between their respective spaces of solutions. To do so, we introduce the set-policy framework which creates a clear link between IL and IRL. As a result, IL and IRL solutions making the best of both worlds are obtained. In addition, it is a unifying framework from which existing IL and IRL algorithms can be derived and which opens the way for IL methods able to deal with the environment's dynamics. Finally, the IRL algorithms derived from the set-policy framework are compared to algorithms belonging to the more common trajectory-matching family. Experiments demonstrate that the set-policy-based algorithms outperform both standard IRL and IL ones and result in more robust solutions.

Mots clés

Learning from Demonstrations Inverse Reinforcement Learning Imitation Learning

Domaines

Machine Learning [stat.ML]

Fichier principal

TNNLS_2016_BPMGOP.pdf (1.57 Mo)

Origine : Fichiers produits par l'(les) auteur(s)

Matthieu GEIST : Connectez-vous pour contacter le contributeur

https://hal.science/hal-01629654

Soumis le : lundi 6 novembre 2017-16:15:54

Dernière modification le : jeudi 11 avril 2024-13:10:04

Dates et versions

hal-01629654 , version 1 (06-11-2017)

Identifiants

HAL Id : hal-01629654 , version 1
DOI : 10.1109/TNNLS.2016.2543000

Citer

Bilal Piot, Matthieu Geist, Olivier Pietquin. Bridging the Gap Between Imitation Learning and Inverse Reinforcement Learning. IEEE Transactions on Neural Networks and Learning Systems, 2017, 28 (8), pp.1814 - 1826. ⟨10.1109/TNNLS.2016.2543000⟩. ⟨hal-01629654⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

INSU CNRS INRIA CENTRALESUPELEC MALIS UMI-GTL CRISTAL UNIV-LORRAINE INRIA2 CRISTAL-SEQUEL LIEC-UL OTELO-UL UNIV-LILLE

620 Consultations

876 Téléchargements

Bridging the Gap Between Imitation Learning and Inverse Reinforcement Learning

Résumé

Mots clés

Domaines

Dates et versions

Identifiants

Citer

Exporter

Collections

Altmetric

Partager