Adaptor Grammars for the Linguist: Word Segmentation Experiments for Very Low-Resource Languages

Abstract : Computational Language Documentation attempts to make the most recent research in speech and language technologies available to linguists working on language preservation and documentation. In this paper, we pursue two main goals along these lines. The first is to improve upon a strong baseline for the unsupervised word discovery task on two very low-resource Bantu languages, taking advantage of the expertise of linguists on these particular languages. The second consists in exploring the Adaptor Grammar framework as a decision and prediction tool for linguists studying a new language. We experiment 162 grammar configurations for each language and show that using Adaptor Grammars for word segmentation enables us to test hypotheses about a language. Specializing a generic grammar with language specific knowledge leads to great improvements for the word discovery task, ultimately achieving a leap of about 30% token F-score from the results of a strong baseline.
Type de document :
Communication dans un congrès
Workshop on Computational Research in Phonetics, Phonology, and Morphology, Oct 2018, Bruxelles, Belgium. pp.32 - 42, 〈10.18653/v1/P17〉
Liste complète des métadonnées

https://hal.archives-ouvertes.fr/hal-01910757
Contributeur : Limsi Publications <>
Soumis le : jeudi 1 novembre 2018 - 21:48:07
Dernière modification le : mardi 12 février 2019 - 01:30:30
Document(s) archivé(s) le : samedi 2 février 2019 - 14:06:27

Fichier

W18-5804.pdf
Fichiers produits par l'(les) auteur(s)

Identifiants

Citation

Pierre Godard, Laurent Besacier, François Yvon, Martine Adda-Decker, Gilles Adda, et al.. Adaptor Grammars for the Linguist: Word Segmentation Experiments for Very Low-Resource Languages. Workshop on Computational Research in Phonetics, Phonology, and Morphology, Oct 2018, Bruxelles, Belgium. pp.32 - 42, 〈10.18653/v1/P17〉. 〈hal-01910757〉

Partager

Métriques

Consultations de la notice

36

Téléchargements de fichiers

17