Preliminary Experiments on Unsupervised Word Discovery in Mboshi - Archive ouverte HAL Accéder directement au contenu
Communication Dans Un Congrès Année : 2016

Preliminary Experiments on Unsupervised Word Discovery in Mboshi

Résumé

The necessity to document thousands of endangered languages encourages the collaboration between linguists and computer scientists in order to provide the documentary linguistics community with the support of automatic processing tools. The French-German ANR-DFG project Breaking the Unwritten Language Barrier (BULB) aims at developing such tools for three mostly unwritten African languages of the Bantu family. For one of them, Mboshi, a language originating from the " Cu-vette " region of the Republic of Congo, we investigate unsuper-vised word discovery techniques from an unsegmented stream of phonemes. We compare different models and algorithms, both monolingual and bilingual, on a new corpus in Mboshi and French, and discuss various ways to represent the data with suitable granularity. An additional French-English corpus allows us to contrast the results obtained on Mboshi and to experiment with more data.
Fichier principal
Vignette du fichier
886_Paper_last.pdf (340.02 Ko) Télécharger le fichier
Origine : Fichiers produits par l'(les) auteur(s)
Loading...

Dates et versions

hal-01350119 , version 1 (29-07-2016)

Identifiants

  • HAL Id : hal-01350119 , version 1

Citer

Pierre Godard, Gilles Adda, Martine Adda-Decker, Alexandre Allauzen, Laurent Besacier, et al.. Preliminary Experiments on Unsupervised Word Discovery in Mboshi. Interspeech 2016, Sep 2016, San-Francisco, United States. ⟨hal-01350119⟩
502 Consultations
396 Téléchargements

Partager

Gmail Facebook X LinkedIn More