Large scale biomedical texts classification: a kNN and an ESA-based approaches

Khadim Dramé; Fleur Mougin; Gayo Diallo

Article Dans Une Revue Journal of Biomedical Semantics Année : 2016

Large scale biomedical texts classification: a kNN and an ESA-based approaches

(1) , (1) , (1)

Khadim Dramé

Fonction : Auteur

Université de Bordeaux

Fleur Mougin

Fonction : Auteur
PersonId : 3272
IdHAL : fleur-mougin
ORCID : 0000-0002-7436-3010
IdRef : 116242337

Université de Bordeaux

Gayo Diallo

Fonction : Auteur
PersonId : 3812
IdHAL : gayo-diallo
ORCID : 0000-0002-9799-9484
IdRef : 112800084

Université de Bordeaux

Résumé

Background With the large and increasing volume of textual data, automated methods for identifying significant topics to classify textual documents have received a growing interest. While many efforts have been made in this direction, it still remains a real challenge. Moreover, the issue is even more complex as full texts are not always freely available. Then, using only partial information to annotate these documents is promising but remains a very ambitious issue. Methods We propose two classification methods: a k-nearest neighbours (kNN)-based approach and an explicit semantic analysis (ESA)-based approach. Although the kNN-based approach is widely used in text classification, it needs to be improved to perform well in this specific classification problem which deals with partial information. Compared to existing kNN-based methods, our method uses classical Machine Learning (ML) algorithms for ranking the labels. Additional features are also investigated in order to improve the classifiers’ performance. In addition, the combination of several learning algorithms with various techniques for fixing the number of relevant topics is performed. On the other hand, ESA seems promising for this classification task as it yielded interesting results in related issues, such as semantic relatedness computation between texts and text classification. Unlike existing works, which use ESA for enriching the bag-of-words approach with additional knowledge-based features, our ESA-based method builds a standalone classifier. Furthermore, we investigate if the results of this method could be useful as a complementary feature of our kNN-based approach. Results Experimental evaluations performed on large standard annotated datasets, provided by the BioASQ organizers, show that the kNN-based method with the Random Forest learning algorithm achieves good performances compared with the current state-of-the-art methods, reaching a competitive f-measure of 0.55% while the ESA-based approach surprisingly yielded reserved results. Conclusions We have proposed simple classification methods suitable to annotate textual documents using only partial information. They are therefore adequate for large multi-label classification and particularly in the biomedical domain. Thus, our work contributes to the extraction of relevant information from unstructured documents in order to facilitate their automated processing. Consequently, it could be used for various purposes, including document indexing, information retrieval, etc.

Domaines

Intelligence artificielle [cs.AI]

Fichier principal

JBS_Paper_KNN.pdf (597.48 Ko)

Origine : Fichiers produits par l'(les) auteur(s)

Gayo Diallo : Connectez-vous pour contacter le contributeur

https://hal.science/hal-01329565

Soumis le : jeudi 9 juin 2016-13:48:22

Dernière modification le : vendredi 23 octobre 2020-16:38:44

Dates et versions

hal-01329565 , version 1 (09-06-2016)

Identifiants

HAL Id : hal-01329565 , version 1
ARXIV : 1606.02976

Citer

Khadim Dramé, Fleur Mougin, Gayo Diallo. Large scale biomedical texts classification: a kNN and an ESA-based approaches. Journal of Biomedical Semantics, 2016. ⟨hal-01329565⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

INSERM LABRI LABRI-MABIOVIS

44 Consultations

112 Téléchargements

Large scale biomedical texts classification: a kNN and an ESA-based approaches

Résumé

Domaines

Dates et versions

Identifiants

Citer

Exporter

Collections

Altmetric

Partager