Factors affecting the effectiveness of biomedical document indexing and retrieval based on terminologies - Archive ouverte HAL Accéder directement au contenu
Article Dans Une Revue Artificial Intelligence in Medicine Année : 2013

Factors affecting the effectiveness of biomedical document indexing and retrieval based on terminologies

Résumé

The aim of this work is to evaluate a set of indexing and retrieval strategies based on the integration of several biomedical terminologies on the available TREC Genomics collections for an ad hoc information retrieval (IR) task.Materials and methodsWe propose a multi-terminology based concept extraction approach to selecting best concepts from free text by means of voting techniques. We instantiate this general approach on four terminologies (MeSH, SNOMED, ICD-10 and GO). We particularly focus on the effect of integrating terminologies into a biomedical IR process, and the utility of using voting techniques for combining the extracted concepts from each document in order to provide a list of unique concepts.ResultsExperimental studies conducted on the TREC Genomics collections show that our multi-terminology IR approach based on voting techniques are statistically significant compared to the baseline. For example, tested on the 2005 TREC Genomics collection, our multi-terminology based IR approach provides an improvement rate of +6.98% in terms of MAP (mean average precision) (p < 0.05) compared to the baseline. In addition, our experimental results show that document expansion using preferred terms in combination with query expansion using terms from top ranked expanded documents improve the biomedical IR effectiveness.ConclusionWe have evaluated several voting models for combining concepts issued from multiple terminologies. Through this study, we presented many factors affecting the effectiveness of biomedical IR system including term weighting, query expansion, and document expansion models. The appropriate combination of those factors could be useful to improve the IR performance.
Fichier principal
Vignette du fichier
Dinh_12322.pdf (895.72 Ko) Télécharger le fichier
Origine : Fichiers produits par l'(les) auteur(s)

Dates et versions

hal-01123496 , version 1 (05-03-2015)

Identifiants

Citer

Duy Dinh, Lynda Tamine, Fatiha Boubekeur. Factors affecting the effectiveness of biomedical document indexing and retrieval based on terminologies. Artificial Intelligence in Medicine, 2013, vol. 57 (n° 2), pp. 155-167. ⟨10.1016/j.artmed.2012.08.006⟩. ⟨hal-01123496⟩
85 Consultations
196 Téléchargements

Altmetric

Partager

Gmail Facebook X LinkedIn More