An unsupervised classification process for large datasets using web reasoning - Archive ouverte HAL Accéder directement au contenu
Communication Dans Un Congrès Année : 2016

An unsupervised classification process for large datasets using web reasoning

Résumé

Determining valuable data among large volumes of data is one of the main challenges in Big Data. We aim to extract knowledge from these sources using a Hierarchical Multi-Label Classification process called Semantic HMC. This process automatically learns a label hierarchy and classifies items from very large data sources. Five steps compose the Semantic HMC process: Indexation, Vectorization, Hierarchization, Resolution and Realization. The first three steps construct automatically the label hierarchy from data sources. The last two steps classify new items according to the label hierarchy. This paper focuses in the last two steps and presents a new highly scalable process to classify items from huge sets of unstructured text by using ontologies and rule-based reasoning. The process is implemented in a scalable and distributed platform to process Big Data and some results are discussed.
Fichier non déposé

Dates et versions

hal-01418148 , version 1 (16-12-2016)

Identifiants

Citer

Rafael Peixoto, Hassan Thomas, Christophe Cruz, Aurélie Bertaux, Nuno Silva. An unsupervised classification process for large datasets using web reasoning. International Workshop on Semantic Big Data, Jun 2016, New York, United States. pp.9:1--9:6, ⟨10.1145/2928294.2928301⟩. ⟨hal-01418148⟩
209 Consultations
0 Téléchargements

Altmetric

Partager

Gmail Facebook X LinkedIn More