What Makes a Speaker Recognizable in TV Broadcast? Going Beyond Speaker Identification Error Rate - Archive ouverte HAL Accéder directement au contenu
Communication Dans Un Congrès Année : 2015

What Makes a Speaker Recognizable in TV Broadcast? Going Beyond Speaker Identification Error Rate

Résumé

Speaker identification approaches for TV broadcast are usually evaluated and compared based on global error rates derived from the overall duration of missed detection, false alarm and confusion. Based on the analysis of the output of the systems submitted to the final round of the French evaluation campaign REPERE, this paper highlights the fact that these average met-rics lead to the incorrect intuition that current state-of-the-art algorithms partially recognize all speakers. Setting aside incorrect diarization and adverse acoustic conditions, we show that their performance is in fact essentially bi-modal: in a given show, either all speech turns of a speaker are correctly identified or none of them are. We then proceed with trying to understand and explain this behavior, through perfomance prediction experiments. These experiments show that the most discriminant speaker characteristics are – first – their total speech duration in the current show and – then only – the amount of training data available to build their acoustic model.
Fichier principal
Vignette du fichier
Charlet2015.pdf (519.81 Ko) Télécharger le fichier
Origine : Fichiers éditeurs autorisés sur une archive ouverte
Loading...

Dates et versions

hal-01433205 , version 1 (06-04-2017)

Identifiants

  • HAL Id : hal-01433205 , version 1

Citer

Delphine Charlet, Johann Poignant, Hervé Bredin, Corinne Fredouille, Sylvain Meignier. What Makes a Speaker Recognizable in TV Broadcast? Going Beyond Speaker Identification Error Rate. ERRARE Workshop, a satellite event of Interspeech 2015., 2015, Sinaia, Romania. ⟨hal-01433205⟩
220 Consultations
72 Téléchargements

Partager

Gmail Facebook X LinkedIn More