# Exact and Monte Carlo Calculations of Integrated Likelihoods for the Latent Class Model

2 SELECT - Model selection in statistical learning
Inria Saclay - Ile de France, LMO - Laboratoire de Mathématiques d'Orsay
Abstract : The latent class model or multivariate multinomial mixture is a powerful approach for clustering categorical data. This model uses a conditional independence assumption given the latent class to which an object is belonging to represent heterogeneous populations. . In this paper, we exploit the fact that a fully Bayesian analysis with Jeffreys non informative prior distributions does not involve technical difficulty to propose an exact expression of the integrated {\em complete-data} likelihood, which is known as being a meaningful model selection criterion in a clustering perspective. Similarly, a Monte Carlo approximation of the integrated {\em observed-data} likelihood can be obtained in two steps: An exact integration over the parameters is followed by an approximation of the sum over all possible partitions through either a frequentist or a Bayesian importance sampling strategy. Then, the exact and the approximate criteria experimentally compete respectively with their standard asymptotic BIC approximations for choosing the number of mixture components. Numerical experiments on simulated data and a biological example highlight that asymptotic criteria are usually dramatically more conservative than the non asymptotic presented criteria, not only for moderate sample sizes as expected but also for quite large sample sizes. It appears that asymptotic standard criteria could often fail to select some interesting structures present in the data. It is also the opportunity to highlight the deep purpose difference between the integrated {\em complete-data} and the {\em observed-data} likelihoods: The integrated {\em complete-data} likelihood is focussing on a cluster analysis view and favors well separated clusters, implying some robustness against model misspecification, while the integrated {\em observed-data} likelihood is focussing on a density estimation view and is expected to provide a consistent estimation of the distribution of the data.
Keywords :
Document type :
Reports
Domain :

Cited literature [14 references]

https://hal.inria.fr/inria-00310137
Contributor : Gilles Celeux Connect in order to contact the contributor
Submitted on : Thursday, August 7, 2008 - 6:45:12 PM
Last modification on : Sunday, June 26, 2022 - 11:48:28 AM
Long-term archiving on: : Thursday, June 3, 2010 - 6:05:38 PM

### Files

RR6609.pdf
Files produced by the author(s)

### Identifiers

• HAL Id : inria-00310137, version 1

### Citation

Christophe Biernacki, Gilles Celeux, Gérard Govaert. Exact and Monte Carlo Calculations of Integrated Likelihoods for the Latent Class Model. [Research Report] RR-6609, INRIA. 2008, pp.25. ⟨inria-00310137⟩

Record views