A Comparison of External Clustering Evaluation Indices in the Context of Imbalanced Data Sets - Archive ouverte HAL Accéder directement au contenu
Communication Dans Un Congrès Année : 2012

A Comparison of External Clustering Evaluation Indices in the Context of Imbalanced Data Sets

Résumé

For highly imbalanced data sets, almost all the instances are labeled as one class, whereas far fewer examples are labeled as the other classes. In this paper, we present an empirical comparison of seven different clustering evaluation indices when used to assess partitions generated from highly imbalanced data sets. Some of the metrics are based on matching of sets (F-measure), information theory (normalized mutual information and adjusted mutual information), and pair of objects counting (Rand and adjusted Rand indices). We also investigate the BCubed metric, which takes into account the concepts of recall, precision, as well as counting pairs. Furthermore, in order to avoid the class size imbalance effect, we propose a modification to the Rand index, referred to as the normalized class size Rand (NCR) index. In terms of results, apart from NCR, our experiments indicate that all the other analyzed indices are not able to deal properly with the problem of class size imbalance.
Fichier non déposé

Dates et versions

hal-00976355 , version 1 (09-04-2014)

Identifiants

Citer

Marcilio de Souto, Andre Coelho, Katti Faceli, Tieme Sakata, Viviane Bonadia, et al.. A Comparison of External Clustering Evaluation Indices in the Context of Imbalanced Data Sets. SBRN 2012, Oct 2012, Curitiba, Brazil. pp.49-54, ⟨10.1109/SBRN.2012.25⟩. ⟨hal-00976355⟩
14 Consultations
0 Téléchargements

Altmetric

Partager

Gmail Facebook X LinkedIn More