Graph based k-means clustering

An original approach to cluster multi-component data sets is proposed that includes an estimation of the number of clusters. Using Prim's algorithm to construct a minimal spanning tree (MST) we show that, under the assumption that the vertices are approximately distributed according to a spatial homogeneous Poisson process, the number of clusters can be accurately estimated by thresholding the sequence of edge lengths added to the MST by Prim's alorithm. This sequence, called the Prim tra jectory, contains suﬃcient information to determine both the number of clusters and the approximate locations of the cluster centroids. The estimated number of clusters and cluster centroids are used to initialize the generalized Lloyd algorithm, also known as k-means, which circumvents its well known initialization problems. We evaluate the false positive rate of our cluster detection algorithm, using Poisson approximations in Euclidean spaces. Applications of this method in the multi/hyper-spectral imagery domain to a satellite view of Paris and to an image of Mars are also presented.

Mots clés

unsupervised classiﬁcation Data partitioning graph-theoretic methods minimal spanning trees similarity measures information theoretic measures multispectral imaging

Domaines

Machine Learning [stat.ML] Traitement du signal et de l'image [eess.SP] Traitement du signal et de l'image [eess.SP]

Fichier principal

GallMCSH-hal.pdf (1 Mo)

Origine : Fichiers produits par l'(les) auteur(s)

Pierre Comon : Connectez-vous pour contacter le contributeur

https://hal.science/hal-00701886

Soumis le : mercredi 24 octobre 2012-16:22:26

Dernière modification le : jeudi 4 avril 2024-21:23:16

Archivage à long terme le : vendredi 25 janvier 2013-03:50:12

Dates et versions

hal-00701886 , version 1 (27-05-2012)

hal-00701886 , version 2 (24-10-2012)

Identifiants

HAL Id : hal-00701886 , version 2
DOI : 10.1016/j.sigpro.2011.12.009

Citer

Laurent Galluccio, Olivier J.J. Michel, Pierre Comon, Alfred O. Hero. Graph based k-means clustering. Signal Processing, 2012, 92 (9), pp.1970-1984. ⟨10.1016/j.sigpro.2011.12.009⟩. ⟨hal-00701886v2⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

UGA CNRS GIPSA GIPSA-DIS I3S GIPSA-CICS UNIV-COTEDAZUR

448 Consultations

4606 Téléchargements