Word confidence estimation for speech translation

Abstract : Word Confidence Estimation (WCE) for machine transla-tion (MT) or automatic speech recognition (ASR) consists in judging each word in the (MT or ASR) hypothesis as correct or incorrect by tagging it with an appropriate label. In the past, this task has been treated separately in ASR or MT con-texts and we propose here a joint estimation of word confi-dence for a spoken language translation (SLT) task involving both ASR and MT. This research work is possible because we built a specific corpus which is first presented. This cor-pus contains 2643 speech utterances for which a quintuplet containing: ASR output (src-asr), verbatim transcript (src-ref), text translation output (tgt-mt), speech translation out-put (tgt-slt) and post-edition of translation (tgt-pe), is made available. The rest of the paper illustrates how such a corpus (made available to the research community) can be used for evaluating word confidence estimators in ASR, MT or SLT scenarios. WCE for SLT could help rescoring SLT output graphs, improving translators productivity (for translation of lectures or movie subtitling) or it could be useful in interac-tive speech-to-speech translation scenarios. Word confidence estimation (WCE), Spoken Language Translation (SLT), Corpus, Joint features.
Type de document :
Communication dans un congrès
International Workshop on Spoken Language Translation, Dec 2014, Lake Tahoe, United States. 2014
Liste complète des métadonnées

Littérature citée [23 références]  Voir  Masquer  Télécharger

https://hal.archives-ouvertes.fr/hal-01110393
Contributeur : Laurent Besacier <>
Soumis le : mercredi 28 janvier 2015 - 14:30:02
Dernière modification le : jeudi 11 octobre 2018 - 08:48:03
Document(s) archivé(s) le : mercredi 29 avril 2015 - 10:30:48

Fichier

iwslt2014-FINAL.pdf
Fichiers produits par l'(les) auteur(s)

Identifiants

  • HAL Id : hal-01110393, version 1

Collections

Citation

Laurent Besacier, Benjamin Lecouteux, Ngoc-Quang Luong, K Hour, M Hadjsalah. Word confidence estimation for speech translation. International Workshop on Spoken Language Translation, Dec 2014, Lake Tahoe, United States. 2014. 〈hal-01110393〉

Partager

Métriques

Consultations de la notice

315

Téléchargements de fichiers

465