Consistency of Random Forests

Random forests are a learning algorithm proposed by Breiman (2001) which combines several randomized decision trees and aggregates their predictions by averaging. Despite its wide usage and outstanding practical performance, little is known about the mathematical properties of the procedure. This disparity between theory and practice originates in the difficulty to simultaneously analyze both the randomization process and the highly data-dependent tree structure. In the present paper, we take a step forward in forest exploration by proving a consistency result for Breiman's (2001) original algorithm in the context of additive regression models. Our analysis also sheds an interesting light on how random forests can nicely adapt to sparsity in high-dimensional settings.

Mots clés

Random forests Dimension reduction Additive model

Domaines

Statistiques [math.ST] Machine Learning [stat.ML] Théorie [stat.TH]

Fichier principal

article.pdf (425.12 Ko)

Origine : Fichiers produits par l'(les) auteur(s)

Erwan Scornet : Connectez-vous pour contacter le contributeur

https://hal.science/hal-00990008

Soumis le : jeudi 21 mai 2015-21:41:00

Dernière modification le : vendredi 19 avril 2024-16:18:54

Archivage à long terme le : jeudi 5 février 2015-10:15:34

Dates et versions

hal-00990008 , version 1 (12-05-2014)

hal-00990008 , version 2 (31-10-2014)

hal-00990008 , version 3 (21-05-2015)

hal-00990008 , version 4 (07-08-2015)

Identifiants

HAL Id : hal-00990008 , version 3
ARXIV : 1405.2881

Citer

Erwan Scornet, Gérard Biau, Jean-Philippe Vert. Consistency of Random Forests. 2014. ⟨hal-00990008v3⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

ENS-PARIS FNCLCC CURIE

1308 Consultations

522 Téléchargements