Sequence-To-Sequence Voice Conversion using F0 and Time Conditioning and Adversarial Learning

Frederik Bous; Laurent Benaroya; Nicolas Obin; Axel Roebel

Pré-Publication, Document De Travail Année : 2021

Sequence-To-Sequence Voice Conversion using F0 and Time Conditioning and Adversarial Learning

(1) , (1) , (1) , (1)

Frederik Bous

Fonction : Auteur

Analyse et synthèse sonores [Paris]

Laurent Benaroya

Fonction : Auteur

Analyse et synthèse sonores [Paris]

Nicolas Obin

Fonction : Auteur

Analyse et synthèse sonores [Paris]

Axel Roebel

Fonction : Auteur
PersonId : 4527
IdHAL : axel-roebel
ORCID : 0000-0001-6136-4391
IdRef : 227186079

Analyse et synthèse sonores [Paris]

Résumé

This paper presents a sequence-to-sequence voice conversion (S2S-VC) algorithm which allows to preserve some aspects of the source speaker during conversion, typically its prosody, which is useful in many real-life application of voice conversion. In S2S-VC, the decoder is usually conditioned on linguistic and speaker embeddings only, with the consequence that only the linguistic content is actually preserved during conversion. In the proposed S2S-VC architecture, the decoder is conditioned explicitly on the desired F0 sequence so that the converted speech has the same F0 as the one of the source speaker, or any F0 defined arbitrarily. Moreover, an adversarial module is further employed so that the S2S-VC is not only optimized on the available true speech samples, but can also take efficiently advantage of the converted speech samples that can be produced by using various conditioning such as speaker identity, F0, or timing.

Domaines

Son [cs.SD] Traitement du signal et de l'image [eess.SP]

Axel Roebel : Connectez-vous pour contacter le contributeur

https://hal.science/hal-03569597

Soumis le : dimanche 13 février 2022-00:19:43

Dernière modification le : samedi 7 octobre 2023-21:36:22

Dates et versions

hal-03569597 , version 1 (13-02-2022)

Identifiants

HAL Id : hal-03569597 , version 1
ARXIV : 2110.03744

Citer

Frederik Bous, Laurent Benaroya, Nicolas Obin, Axel Roebel. Sequence-To-Sequence Voice Conversion using F0 and Time Conditioning and Adversarial Learning. 2021. ⟨hal-03569597⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

CNRS IRCAM STMS SORBONNE-UNIVERSITE SU-SCIENCES

39 Consultations

0 Téléchargements

Sequence-To-Sequence Voice Conversion using F0 and Time Conditioning and Adversarial Learning

Résumé

Domaines

Dates et versions

Identifiants

Citer

Exporter

Collections

Altmetric

Partager