déposer
version française rss feed
HAL : hal-00558145, version 2

Fiche détaillée  Récupérer au format
Versions disponibles :
A fast and recursive algorithm for clustering large datasets with $k$-medians
Hervé Cardot 1, Peggy Cenac 1, Jean-Marie Monnez 2, 3
BIGS Collaboration(s)
(21/01/2011)

Clustering with fast algorithms large samples of high dimensional data is an important challenge in computational statistics. A new class of recursive stochastic gradient algorithms designed for the $k$-medians loss criterion is proposed. By their recursive nature, these algorithms are very fast and are well adapted to deal with large samples of data that are allowed to arrive sequentially. It is proved that the stochastic gradient algorithm converges almost surely to the set of stationary points of the underlying loss criterion. A particular attention is paid to the averaged versions which are known to have better performances. A data-driven procedure that permits a fully automatic selection of the value of the descent step is also proposed. The performance of the averaged sequential estimator is compared on a simulation study, both in terms of computation speed and accuracy of the estimations, with more classical partitioning techniques such as $k$-means, trimmed $k$-means and PAM (partitioning around medoids). Finally, this new online clustering technique is illustrated on determining television audience profiles with a sample of more than 5000 individual television audiences measured every minute over a period of 24 hours.
1 :  Institut de Mathématiques de Bourgogne (IMB)
CNRS : UMR5584 – Université de Bourgogne
2 :  Institut Elie Cartan Nancy (IECN)
CNRS : UMR7502 – INRIA – Université Henri Poincaré - Nancy I – Université Nancy II – Institut National Polytechnique de Lorraine (INPL)
3 :  BIGS (INRIA Lorraine / IECN)
INRIA – CNRS : UMR7502
Probabilités et statistiques
Mathématiques/Statistiques

Statistiques/Théorie
Averaging – High dimensional data – Partitioning around medoids – Recursive estimator – Stochastic approximation
Liste des fichiers attachés à ce document : 
PDF
k-medianv5.pdf(1.5 MB)

tous les articles de la base du CCSd...
tous les articles de la base du CCSd...
tous les articles de la base du CCSd...
tous les articles de la base du CCSd...
tous les articles de la base du CCSd...
tous les articles de la base du CCSd...
tous les articles de la base du CCSd...
tous les articles de la base du CCSd...
tous les articles de la base du CCSd...
tous les articles de la base du CCSd...
tous les articles de la base du CCSd...
tous les articles de la base du CCSd...
tous les articles de la base du CCSd...
tous les articles de la base du CCSd...
tous les articles de la base du CCSd...
tous les articles de la base du CCSd...
tous les articles de la base du CCSd...
tous les articles de la base du CCSd...
tous les articles de la base du CCSd...
tous les articles de la base du CCSd...