Jira: a Kurdish Speech Recognition System Designing and Building Speech Corpus and Pronunciation Lexicon

Hadi Veisi; Hawre Hosseini; Mohammad Mohammadamini; Wirya Fathy; Aso Mahmudi

Pré-Publication, Document De Travail Année : 2021

Jira: a Kurdish Speech Recognition System Designing and Building Speech Corpus and Pronunciation Lexicon

(1) , (2) , (3) , (1) , (1)

1
2
3

Hadi Veisi

Fonction : Auteur
PersonId : 1090899

University of Tehran

Hawre Hosseini

Fonction : Auteur
PersonId : 1090900

Ryerson University [Toronto]

Mohammad Mohammadamini

Fonction : Auteur
PersonId : 1070002

Laboratoire Informatique d'Avignon

Wirya Fathy

Fonction : Auteur
PersonId : 1090901

University of Tehran

Aso Mahmudi

Fonction : Auteur
PersonId : 1090902

University of Tehran

Résumé

In this paper, we introduce the first large vocabulary speech recognition system (LVSR) for the Central Kurdish language, named Jira. The Kurdish language is an Indo-European language spoken by more than 30 million people in several countries, but due to the lack of speech and text resources, there is no speech recognition system for this language. To fill this gap, we introduce the first speech corpus and pronunciation lexicon for the Kurdish language. Regarding speech corpus, we designed a sentence collection in which the ratio of di-phones in the collection resembles the real data of the Central Kurdish language. The designed sentences are uttered by 576 speakers in a controlled environment with noise-free microphones (called AsoSoft Speech-Office) and in Telegram social network environment using mobile phones (denoted as AsoSoft Speech-Crowdsourcing), resulted in 43.68 hours of speech. Besides, a test set including 11 different document topics is designed and recorded in two corresponding speech conditions (i.e., Office and Crowdsourcing). Furthermore, a 60K pronunciation lexicon is prepared in this research in which we faced several challenges and proposed solutions for them. The Kurdish language has several dialects and sub-dialects that results in many lexical variations. Our methods for script standardization of lexical variations and automatic pronunciation of the lexicon tokens are presented in detail. To setup the recognition engine, we used the Kaldi toolkit. A statistical tri-gram language model that is extracted from the AsoSoft text corpus is used in the system. Several standard recipes including HMM-based models (i.e., mono, tri1, tr2, tri2, tri3), SGMM, and DNN methods are used to generate the acoustic model. These methods are trained with AsoSoft Speech-Office and AsoSoft Speech-Crowdsourcing and a combination of them. The best performance achieved by the SGMM acoustic model which results in 13.9% of the average word error rate (on different document topics) and 4.9% for the general topic.

Mots clés

Kurdish Language Speech Corpus Pronunciation Lexicon Speech Recognition Jira

Domaines

Intelligence artificielle [cs.AI] Linguistique

Fichier principal

Jira a Kurdish Speech Recognition System.pdf (575.03 Ko)

Origine : Fichiers produits par l'(les) auteur(s)

Mohammad Mohammadamini : Connectez-vous pour contacter le contributeur

https://hal.science/hal-03140680

Soumis le : samedi 13 février 2021-14:35:48

Dernière modification le : lundi 15 février 2021-10:40:40

Archivage à long terme le : vendredi 14 mai 2021-18:18:11

Dates et versions

hal-03140680 , version 1 (13-02-2021)

Identifiants

HAL Id : hal-03140680 , version 1
ARXIV : 2102.07412

Citer

Hadi Veisi, Hawre Hosseini, Mohammad Mohammadamini, Wirya Fathy, Aso Mahmudi. Jira: a Kurdish Speech Recognition System Designing and Building Speech Corpus and Pronunciation Lexicon. 2021. ⟨hal-03140680⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

UNIV-AVIGNON LIA

148 Consultations

388 Téléchargements

Jira: a Kurdish Speech Recognition System Designing and Building Speech Corpus and Pronunciation Lexicon

Résumé

Mots clés

Domaines

Dates et versions

Identifiants

Citer

Exporter

Collections

Altmetric

Partager