Learning from narrated instruction videos - Archive ouverte HAL Accéder directement au contenu
Article Dans Une Revue IEEE Transactions on Pattern Analysis and Machine Intelligence Année : 2017

Learning from narrated instruction videos

Résumé

Automatic assistants could guide a person or a robot in performing new tasks, such as changing a car tire or repotting a plant. Creating such assistants, however, is non-trivial and requires understanding of visual and verbal content of a video. Towards this goal, we here address the problem of automatically learning the main steps of a task from a set of narrated instruction videos. We develop a new unsupervised learning approach that takes advantage of the complementary nature of the input video and the associated narration. The method sequentially clusters textual and visual representations of a task, where the two clustering problems are linked by joint constraints to obtain a single coherent sequence of steps in both modalities. To evaluate our method, we collect and annotate a new challenging dataset of real-world instruction videos from the Internet. The dataset contains videos for five different tasks with complex interactions between people and objects, captured in a variety of indoor and outdoor settings. We experimentally demonstrate that the proposed method can automatically discover, learn and localize the main steps of a task in input videos.
Fichier principal
Vignette du fichier
pami2016alayrac.pdf (6.88 Mo) Télécharger le fichier
Origine : Fichiers produits par l'(les) auteur(s)
Loading...

Dates et versions

hal-01580630 , version 1 (01-09-2017)

Identifiants

  • HAL Id : hal-01580630 , version 1

Citer

Jean-Baptiste Alayrac, Piotr Bojanowski, Nishant Agrawal, Josef Sivic, Ivan Laptev, et al.. Learning from narrated instruction videos. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017, XX. ⟨hal-01580630⟩
525 Consultations
613 Téléchargements

Partager

Gmail Facebook X LinkedIn More