Learning cross-lingual phonological and orthagraphic adaptations: a case study in improving neural machine translation between low-resource languages - Archive ouverte HAL Accéder directement au contenu
Article Dans Une Revue Journal of Language Modelling Année : 2019

Learning cross-lingual phonological and orthagraphic adaptations: a case study in improving neural machine translation between low-resource languages

Saurav Jha
  • Fonction : Auteur
Akhilesh Sudhakar
  • Fonction : Auteur

Résumé

Out-of-vocabulary (OOV) words can pose serious challenges for machine translation (MT) tasks, and in particular, for low-resource language (LRL) pairs, i.e., language pairs for which few or no parallel corpora exist. Our work adapts variants of seq2seq models to perform transduction of such words from Hindi to Bhojpuri (an LRL instance), learning from a set of cognate pairs built from a bilingual dictionary of Hindi - Bhojpuri words. We demonstrate that our models can be effectively used for language pairs that have limited parallel corpora; our models work at the character level to grasp phonetic and orthographic similarities across multiple types of word adaptations, whether synchronic or diachronic, loan words or cognates. We describe the training aspects of several character level NMT systems that we adapted to this task and characterize their typical errors. Our method improves BLEU score by 6.3 on the Hindi-to-Bhojpuri translation task. Further, we show that such transductions can generalize well to other languages by applying it successfully to Hindi - Bangla cognate pairs. Our work can be seen as an important step in the process of: (i) resolving the OOV words problem arising in MT tasks; (ii) creating effective parallel corpora for resource constrained languages; and (iii) leveraging the enhanced semantic knowledge captured by word-level embeddings to perform character-level tasks.
Fichier principal
Vignette du fichier
word-transduction-jlm-2018-camera-ready-v6.pdf (516.75 Ko) Télécharger le fichier
Origine : Fichiers éditeurs autorisés sur une archive ouverte

Dates et versions

hal-03036776 , version 1 (02-12-2020)

Licence

Paternité

Identifiants

  • HAL Id : hal-03036776 , version 1

Citer

Saurav Jha, Akhilesh Sudhakar, Anil Kumar Singh. Learning cross-lingual phonological and orthagraphic adaptations: a case study in improving neural machine translation between low-resource languages. Journal of Language Modelling, 2019, Special issue on finite-state methods in natural language processing and mathematics of language, 7 (2), pp.101-142. ⟨hal-03036776⟩

Collections

TDS-MACS
13 Consultations
38 Téléchargements

Partager

Gmail Facebook X LinkedIn More