A Trainable Optimal Transport Embedding for Feature Aggregation and its Relationship to Attention

Grégoire Mialon; Dexiong Chen; Alexandre d'Aspremont; Julien Mairal

Communication Dans Un Congrès Année : 2021

A Trainable Optimal Transport Embedding for Feature Aggregation and its Relationship to Attention

(1, 2) , (1) , (2) , (1)

1
2

Grégoire Mialon

Fonction : Auteur
PersonId : 1036976

Apprentissage de modèles à partir de données massives

Statistical Machine Learning and Parsimony

Dexiong Chen

Fonction : Auteur
PersonId : 1047920

Apprentissage de modèles à partir de données massives

Alexandre d'Aspremont

Fonction : Auteur
PersonId : 10163
IdHAL : aspremon
ORCID : 0000-0003-3851-216X
IdRef : 157968219

Statistical Machine Learning and Parsimony

Julien Mairal

Fonction : Auteur
PersonId : 1034832
ORCID : 0000-0001-6991-2110
IdRef : 152125256

Apprentissage de modèles à partir de données massives

Résumé

We address the problem of learning on sets of features, motivated by the need of performing pooling operations in long biological sequences of varying sizes, with long-range dependencies, and possibly few labeled data. To address this challenging task, we introduce a parametrized representation of fixed size, which embeds and then aggregates elements from a given input set according to the optimal transport plan between the set and a trainable reference. Our approach scales to large datasets and allows end-to-end training of the reference, while also providing a simple unsupervised learning mechanism with small computational cost. Our aggregation technique admits two useful interpretations: it may be seen as a mechanism related to attention layers in neural networks, or it may be seen as a scalable surrogate of a classical optimal transport-based kernel. We experimentally demonstrate the effectiveness of our approach on biological sequences, achieving state-of-the-art results for protein fold recognition and detection of chromatin profiles tasks, and, as a proof of concept, we show promising results for processing natural language sequences. We provide an open-source implementation of our embedding that can be used alone or as a module in larger learning models at https://github.com/claying/OTK.

Domaines

Machine Learning [stat.ML] Apprentissage [cs.LG]

Fichier principal

main_iclr.pdf (1.21 Mo)

Origine : Fichiers produits par l'(les) auteur(s)

Grégoire Mialon : Connectez-vous pour contacter le contributeur

https://hal.science/hal-02883436

Soumis le : mardi 9 février 2021-15:05:23

Dernière modification le : samedi 27 avril 2024-03:10:50

Dates et versions

hal-02883436 , version 1 (29-06-2020)

hal-02883436 , version 2 (05-10-2020)

hal-02883436 , version 3 (09-02-2021)

Identifiants

HAL Id : hal-02883436 , version 3

Citer

Grégoire Mialon, Dexiong Chen, Alexandre d'Aspremont, Julien Mairal. A Trainable Optimal Transport Embedding for Feature Aggregation and its Relationship to Attention. ICLR 2021 - The Ninth International Conference on Learning Representations, May 2021, Virtual, France. ⟨hal-02883436v3⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

ENS-PARIS UNIV-RENNES1 UGA CNRS INRIA IRISA INSMI LJK LJK_GI INRIA2 LJK-GI-THOTH PSL UR1-MATH-STIC UR1-UFR-ISTIC UNIV-RENNES MIAI ANR PRAIRIE-IA UR1-MATH-NUM

6463 Consultations

697 Téléchargements

A Trainable Optimal Transport Embedding for Feature Aggregation and its Relationship to Attention

Résumé

Domaines

Dates et versions

Identifiants

Citer

Exporter

Collections

Partager