Wavelet Scattering Transform and CNN for Closed Set Speaker Identification

Wajdi Ghezaiel; Luc Brun; Olivier Lézoray

Communication Dans Un Congrès Année : 2020

Wavelet Scattering Transform and CNN for Closed Set Speaker Identification

(1) , (2) , (2)

1
2

Wajdi Ghezaiel

Fonction : Auteur
PersonId : 1078474

Fédération Normande de Recherche en Sciences et Technologies de l’Information et de la Communication

Luc Brun

Fonction : Auteur
PersonId : 936962

Equipe Image - Laboratoire GREYC - UMR6072

Olivier Lézoray

Fonction : Auteur
PersonId : 230
IdHAL : olivier-lezoray
ORCID : 0000-0003-0540-543X
IdRef : 17227687X

Equipe Image - Laboratoire GREYC - UMR6072

Résumé

In real world applications, the performances of speaker identification systems degrade due to the reduction of both the amount and the quality of speech utterance. For that particular purpose, we propose a speaker identification system where short utterances with few training examples are used for person identification. Therefore, only a very small amount of data involving a sentence of 2-4 seconds is used. To achieve this, we propose a novel raw waveform end-to-end convolutional neural network (CNN) for text-independent speaker identification. We use wavelet scattering transform as a fixed initialization of the first layers of a CNN network, and learn the remaining layers in a supervised manner. The conducted experiments show that our hybrid architecture combining wavelet scattering transform and CNN can successfully perform efficient feature extraction for a speaker identification, even with a small number of short duration training samples.

Mots clés

Speaker identification short utterances wavelet scattering transform convolutional neural network hybrid net- work

Domaines

Traitement des images [eess.IV]

Fichier principal

MMSP2020.pdf (245.88 Ko)

Origine : Fichiers produits par l'(les) auteur(s)

Olivier Lezoray : Connectez-vous pour contacter le contributeur

https://hal.science/hal-02955532

Soumis le : vendredi 2 octobre 2020-08:16:48

Dernière modification le : mercredi 20 mars 2024-16:20:04

Archivage à long terme le : lundi 4 janvier 2021-08:49:07

Dates et versions

hal-02955532 , version 1 (02-10-2020)

Identifiants

HAL Id : hal-02955532 , version 1

Citer

Wajdi Ghezaiel, Luc Brun, Olivier Lézoray. Wavelet Scattering Transform and CNN for Closed Set Speaker Identification. International Workshop on Multimedia Signal Processing (MMSP), Sep 2020, Tampere (Virtual conference), Finland. ⟨hal-02955532⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

CNRS INSA-ROUEN GREYC GREYC-IMAGE COMUE-NORMANDIE UNIROUEN ENSICAEN UNILEHAVRE UNICAEN INSA-GROUPE

138 Consultations

1424 Téléchargements

Wavelet Scattering Transform and CNN for Closed Set Speaker Identification

Résumé

Mots clés

Domaines

Dates et versions

Identifiants

Citer

Exporter

Collections

Partager