Voting Classifier vs Deep learning method in Arabic Dialect Identification

Dhaou Ghoul; Gaël Lejeune

Communication Dans Un Congrès Année : 2020

Voting Classifier vs Deep learning method in Arabic Dialect Identification

(1, 2) , (3, 2)

1
2
3

Dhaou Ghoul

Fonction : Auteur
PersonId : 1072388

Sens, Texte, Informatique, Histoire

Équipe Linguistique computationnelle

Gaël Lejeune

Fonction : Auteur
PersonId : 734695
IdHAL : gael-lejeune
ORCID : 0000-0002-4795-2362
IdRef : 182283054

Sens, Texte, Informatique, Histoire

Équipe Linguistique computationnelle

Résumé

In this paper, we present three methods developed by the SORBONNE Team for the NADI shared task on Arabic Dialect Identification for tweets. The first and the second method use respectively a machine learning model based on a Voting Classifier with words and character level features and a deep learning model at the word level. The third method uses only character-level features. We explored different text representation such as TF-IDF (first model) and word embeddings (second model). The Voting Classifier was the most powerful prediction model, achieving the best macro-average F1 score of 18.8% and an accuracy of 36.54% on the official test. Our model ranked 9 on the challenge and in conclusion we propose some ideas to improve its results.

Domaines

Traitement du texte et du document

Fichier principal

coling2020.pdf (220.85 Ko)

Origine : Fichiers produits par l'(les) auteur(s)

Gaël Lejeune : Connectez-vous pour contacter le contributeur

https://hal.science/hal-03089957

Soumis le : mardi 29 décembre 2020-07:26:47

Dernière modification le : samedi 7 octobre 2023-21:36:24

Archivage à long terme le : mardi 30 mars 2021-18:05:59

Dates et versions

hal-03089957 , version 1 (29-12-2020)

Identifiants

HAL Id : hal-03089957 , version 1

Citer

Dhaou Ghoul, Gaël Lejeune. Voting Classifier vs Deep learning method in Arabic Dialect Identification. : Proceedings of the Fifth Arabic Natural Language Processing Workshop, COLING 2020, Dec 2020, Barcelone, Spain. ⟨hal-03089957⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

SORBONNE-UNIVERSITE STIH SU-LETTRES

58 Consultations

53 Téléchargements

Voting Classifier vs Deep learning method in Arabic Dialect Identification

Résumé

Domaines

Dates et versions

Identifiants

Citer

Exporter

Collections

Partager