Robust Discriminative Clustering with Sparse Regularizers

Nicolas Flammarion 1, 2 Balamurugan Palaniappan 1, 2 Francis Bach 1, 2
1 SIERRA - Statistical Machine Learning and Parsimony
DI-ENS - Département d'informatique de l'École normale supérieure, CNRS - Centre National de la Recherche Scientifique, Inria de Paris
Abstract : Clustering high-dimensional data often requires some form of dimensionality reduction, where clustered variables are separated from " noise-looking " variables. We cast this problem as finding a low-dimensional projection of the data which is well-clustered. This yields a one-dimensional projection in the simplest situation with two clusters, and extends naturally to a multi-label scenario for more than two clusters. In this paper, (a) we first show that this joint clustering and dimension reduction formulation is equivalent to previously proposed discriminative clustering frameworks, thus leading to convex relaxations of the problem; (b) we propose a novel sparse extension, which is still cast as a convex relaxation and allows estimation in higher dimensions; (c) we propose a natural extension for the multi-label scenario; (d) we provide a new theoretical analysis of the performance of these formulations with a simple probabilistic model, leading to scalings over the form d = O(√ n) for the affine invariant case and d = O(n) for the sparse case, where n is the number of examples and d the ambient dimension; and finally, (e) we propose an efficient iterative algorithm with running-time complexity proportional to O(nd 2), improving on earlier algorithms which had quadratic complexity in the number of examples.
Type de document :
Article dans une revue
Journal of Machine Learning Research (JMLR), 2017, 18 (80), pp.1-50
Liste complète des métadonnées

Littérature citée [31 références]  Voir  Masquer  Télécharger

https://hal.archives-ouvertes.fr/hal-01357666
Contributeur : Nicolas Flammarion <>
Soumis le : mardi 30 août 2016 - 11:44:14
Dernière modification le : jeudi 26 avril 2018 - 10:29:11

Fichier

RobDis_Hal.pdf
Fichiers produits par l'(les) auteur(s)

Identifiants

  • HAL Id : hal-01357666, version 1
  • ARXIV : 1608.08052

Collections

Citation

Nicolas Flammarion, Balamurugan Palaniappan, Francis Bach. Robust Discriminative Clustering with Sparse Regularizers. Journal of Machine Learning Research (JMLR), 2017, 18 (80), pp.1-50. 〈hal-01357666〉

Partager

Métriques

Consultations de la notice

296

Téléchargements de fichiers

141