A bagging SVM to learn from positive and unlabeled examples

Fantine Mordelet; Jean-Philippe Vert

Pré-Publication, Document De Travail Année : 2010

A bagging SVM to learn from positive and unlabeled examples

(1, 2) , (1, 2)

1
2

Fantine Mordelet

Fonction : Auteur
PersonId : 847190

Cancer et génome: Bioinformatique, biostatistiques et épidémiologie d'un système complexe

Centre de Bioinformatique

Jean-Philippe Vert

Fonction : Auteur
PersonId : 10276
IdHAL : jean-philippe-vert
ORCID : 0000-0001-9510-8441
IdRef : 122407385

Cancer et génome: Bioinformatique, biostatistiques et épidémiologie d'un système complexe

Centre de Bioinformatique

Résumé

We consider the problem of learning a binary classifier from a training set of positive and unlabeled examples, both in the inductive and in the transductive setting. This problem, often referred to as \emph{PU learning}, differs from the standard supervised classification problem by the lack of negative examples in the training set. It corresponds to an ubiquitous situation in many applications such as information retrieval or gene ranking, when we have identified a set of data of interest sharing a particular property, and we wish to automatically retrieve additional data sharing the same property among a large and easily available pool of unlabeled data. We propose a conceptually simple method, akin to bagging, to approach both inductive and transductive PU learning problems, by converting them into series of supervised binary classification problems discriminating the known positive examples from random subsamples of the unlabeled set. We empirically demonstrate the relevance of the method on simulated and real data, where it performs at least as well as existing methods while being faster.

Mots clés

PU learning Bagging SVM

Domaines

Machine Learning [stat.ML]

Fichier principal

PUL-techreport.pdf (231.99 Ko)

Origine : Fichiers produits par l'(les) auteur(s)

Fantine Mordelet : Connectez-vous pour contacter le contributeur

https://hal.science/hal-00523336

Soumis le : lundi 4 octobre 2010-19:10:33

Dernière modification le : vendredi 19 avril 2024-16:18:56

Archivage à long terme le : mercredi 5 janvier 2011-03:13:19

Dates et versions

hal-00523336 , version 1 (04-10-2010)

Identifiants

HAL Id : hal-00523336 , version 1
ARXIV : 1010.0772

Citer

Fantine Mordelet, Jean-Philippe Vert. A bagging SVM to learn from positive and unlabeled examples. 2010. ⟨hal-00523336⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

INSTITUT-TELECOM ENSMP ENSMP_CBIO PARISTECH FNCLCC CURIE PSL ENSMP_DR

835 Consultations

612 Téléchargements

A bagging SVM to learn from positive and unlabeled examples

Résumé

Mots clés

Domaines

Dates et versions

Identifiants

Citer

Exporter

Collections

Altmetric

Partager