Beyond Voice Identity Conversion: Manipulating Voice Attributes by Adversarial Learning of Structured Disentangled Representations

Laurent Benaroya; Nicolas Obin; Axel Roebel

Pré-Publication, Document De Travail Année : 2021

Beyond Voice Identity Conversion: Manipulating Voice Attributes by Adversarial Learning of Structured Disentangled Representations

(1) , (1) , (1)

Laurent Benaroya

Fonction : Auteur

Analyse et synthèse sonores [Paris]

Nicolas Obin

Fonction : Auteur

Analyse et synthèse sonores [Paris]

Axel Roebel

Fonction : Auteur
PersonId : 4527
IdHAL : axel-roebel
ORCID : 0000-0001-6136-4391
IdRef : 227186079

Analyse et synthèse sonores [Paris]

Résumé

Voice conversion (VC) consists of digitally altering the voice of an individual to manipulate part of its content, primarily its identity, while maintaining the rest unchanged. Research in neural VC has accomplished considerable breakthroughs with the capacity to falsify a voice identity using a small amount of data with a highly realistic rendering. This paper goes beyond voice identity and presents a neural architecture that allows the manipulation of voice attributes (e.g., gender and age). Leveraging the latest advances on adversarial learning of structured speech representation, a novel structured neural network is proposed in which multiple auto-encoders are used to encode speech as a set of idealistically independent linguistic and extra-linguistic representations, which are learned adversariarly and can be manipulated during VC. Moreover, the proposed architecture is time-synchronized so that the original voice timing is preserved during conversion which allows lip-sync applications. Applied to voice gender conversion on the real-world VCTK dataset, our proposed architecture can learn successfully gender-independent representation and convert the voice gender with a very high efficiency and naturalness.

Domaines

Son [cs.SD] Traitement du signal et de l'image [eess.SP]

Axel Roebel : Connectez-vous pour contacter le contributeur

https://hal.science/hal-03569608

Soumis le : dimanche 13 février 2022-00:23:16

Dernière modification le : samedi 7 octobre 2023-21:36:22

Dates et versions

hal-03569608 , version 1 (13-02-2022)

Identifiants

HAL Id : hal-03569608 , version 1
ARXIV : 2107.12346

Citer

Laurent Benaroya, Nicolas Obin, Axel Roebel. Beyond Voice Identity Conversion: Manipulating Voice Attributes by Adversarial Learning of Structured Disentangled Representations. 2021. ⟨hal-03569608⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

CNRS IRCAM STMS SORBONNE-UNIVERSITE SU-SCIENCES ANR

37 Consultations

0 Téléchargements

Beyond Voice Identity Conversion: Manipulating Voice Attributes by Adversarial Learning of Structured Disentangled Representations

Résumé

Domaines

Dates et versions

Identifiants

Citer

Exporter

Collections

Altmetric

Partager