Preliminary Experiments on Unsupervised Word Discovery in Mboshi

Abstract : The necessity to document thousands of endangered languages encourages the collaboration between linguists and computer scientists in order to provide the documentary linguistics community with the support of automatic processing tools. The French-German ANR-DFG project Breaking the Unwritten Language Barrier (BULB) aims at developing such tools for three mostly unwritten African languages of the Bantu family. For one of them, Mboshi, a language originating from the " Cu-vette " region of the Republic of Congo, we investigate unsuper-vised word discovery techniques from an unsegmented stream of phonemes. We compare different models and algorithms, both monolingual and bilingual, on a new corpus in Mboshi and French, and discuss various ways to represent the data with suitable granularity. An additional French-English corpus allows us to contrast the results obtained on Mboshi and to experiment with more data.
Document type :
Conference papers
Liste complète des métadonnées

Cited literature [32 references]  Display  Hide  Download

https://hal.archives-ouvertes.fr/hal-01350119
Contributor : Laurent Besacier <>
Submitted on : Friday, July 29, 2016 - 4:33:32 PM
Last modification on : Thursday, April 4, 2019 - 10:18:05 AM
Document(s) archivé(s) le : Sunday, October 30, 2016 - 10:46:43 AM

File

886_Paper_last.pdf
Files produced by the author(s)

Identifiers

  • HAL Id : hal-01350119, version 1

Citation

Pierre Godard, Gilles Adda, Martine Adda-Decker, Alexandre Allauzen, Laurent Besacier, et al.. Preliminary Experiments on Unsupervised Word Discovery in Mboshi. Interspeech 2016, Sep 2016, San-Francisco, United States. ⟨hal-01350119⟩

Share

Metrics

Record views

562

Files downloads

307