Skip to Main content Skip to Navigation
Journal articles

Data-driven calibration of penalties for least-squares regression

Sylvain Arlot 1, 2 Pascal Massart 1, 2
2 SELECT - Model selection in statistical learning
LMO - Laboratoire de Mathématiques d'Orsay, Inria Saclay - Ile de France
Abstract : Penalization procedures often suffer from their dependence on multiplying factors, whose optimal values are either unknown or hard to estimate from the data. We propose a completely data-driven calibration algorithm for this parameter in the least-squares regression framework, without assuming a particular shape for the penalty. Our algorithm relies on the concept of minimal penalty, recently introduced by Birge and Massart (2007) in the context of penalized least squares for Gaussian homoscedastic regression. On the positive side, the minimal penalty can be evaluated from the data themselves, leading to a data-driven estimation of an optimal penalty which can be used in practice; on the negative side, their approach heavily relies on the homoscedastic Gaussian nature of their stochastic framework. The purpose of this paper is twofold: stating a more general heuristics for designing a data-driven penalty (the slope heuristics) and proving that it works for penalized least-squares regression with a random design, even for heteroscedastic non-Gaussian data. For technical reasons, some exact mathematical results will be proved only for regressogram bin-width selection. This is at least a first step towards further results, since the approach and the method that we use are indeed general.
Complete list of metadatas

Cited literature [44 references]  Display  Hide  Download

https://hal.archives-ouvertes.fr/hal-00243116
Contributor : Sylvain Arlot <>
Submitted on : Wednesday, December 17, 2008 - 10:17:04 AM
Last modification on : Friday, November 27, 2020 - 5:50:03 PM
Long-term archiving on: : Friday, September 24, 2010 - 10:36:39 AM

Files

arlot08a.pdf
Files produced by the author(s)

Identifiers

  • HAL Id : hal-00243116, version 4
  • ARXIV : 0802.0837

Collections

Citation

Sylvain Arlot, Pascal Massart. Data-driven calibration of penalties for least-squares regression. Journal of Machine Learning Research, Microtome Publishing, 2009, 10, pp.245-279. ⟨hal-00243116v4⟩

Share

Metrics

Record views

834

Files downloads

343