An unsupervised classification process for large datasets using web reasoning

Abstract : Determining valuable data among large volumes of data is one of the main challenges in Big Data. We aim to extract knowledge from these sources using a Hierarchical Multi-Label Classification process called Semantic HMC. This process automatically learns a label hierarchy and classifies items from very large data sources. Five steps compose the Semantic HMC process: Indexation, Vectorization, Hierarchization, Resolution and Realization. The first three steps construct automatically the label hierarchy from data sources. The last two steps classify new items according to the label hierarchy. This paper focuses in the last two steps and presents a new highly scalable process to classify items from huge sets of unstructured text by using ontologies and rule-based reasoning. The process is implemented in a scalable and distributed platform to process Big Data and some results are discussed.
Liste complète des métadonnées
Contributor : Aurélie Bertaux <>
Submitted on : Friday, December 16, 2016 - 1:31:55 PM
Last modification on : Wednesday, September 12, 2018 - 1:27:08 AM




Rafael Peixoto, Hassan Thomas, Christophe Cruz, Aurélie Bertaux, Nuno Silva. An unsupervised classification process for large datasets using web reasoning. International Workshop on Semantic Big Data, Jun 2016, New York, United States. pp.9:1--9:6, ⟨10.1145/2928294.2928301⟩. ⟨hal-01418148⟩



Record views