Innovative technologies for under-resourced language documentation: The BULB Project

Abstract : The project Breaking the Unwritten Language Barrier (BULB), which brings together linguists and computer scientists, aims at supporting linguists in documenting unwritten languages. In order to achieve this we will develop tools tailored to the needs of documentary linguists by building upon technology and expertise from the area of natural language processing, most prominently automatic speech recognition and machine translation. As a development and test bed for this we have chosen three less-resourced African languages from the Bantu family: Basaa, Myene and Embosi. Work within the project is divided into three main steps: 1) Collection of a large corpus of speech (100h per language) at a reasonable cost. After initial recording, the data is re-spoken by a reference speaker to enhance the signal quality and orally translated into French. 2) Automatic transcription of the Bantu languages at phoneme level and the French translation at word level. The recognized Bantu phonemes and French words will then be automatically aligned. 3) Tool development. In close cooperation and discussion with the linguists, the speech and language technologists will design and implement tools that will support the linguists in their work, taking into account the linguists' needs and technology's capabilities. The data collection has begun for the three languages. For this we use standard mobile devices and a dedicated software—LIG-AIKUMA, which proposes a range of different speech collection modes (recording, respeaking, translation and elicitation). LIG-AIKUMA 's improved features include a smart generation and handling of speaker metadata as well as respeaking and parallel audio data mapping.
Type de document :
Communication dans un congrès
Workshop CCURL 2016 - Collaboration and Computing for Under-Resourced Languages - LREC, May 2016, Portoroz, Slovenia. CCURL proceedings
Liste complète des métadonnées

Littérature citée [42 références]  Voir  Masquer  Télécharger

https://hal.archives-ouvertes.fr/hal-01350124
Contributeur : Laurent Besacier <>
Soumis le : vendredi 29 juillet 2016 - 16:48:10
Dernière modification le : lundi 17 décembre 2018 - 10:35:29

Fichier

CCURL_BULB_2016.pdf
Fichiers produits par l'(les) auteur(s)

Identifiants

  • HAL Id : hal-01350124, version 1

Citation

Gilles Adda, Martine Adda-Decker, Odette Ambouroue, Laurent Besacier, David Blachon, et al.. Innovative technologies for under-resourced language documentation: The BULB Project. Workshop CCURL 2016 - Collaboration and Computing for Under-Resourced Languages - LREC, May 2016, Portoroz, Slovenia. CCURL proceedings. 〈hal-01350124〉

Partager

Métriques

Consultations de la notice

1057

Téléchargements de fichiers

305