From text saliency to linguistic objects: learning linguistic interpretable markers with a multi-channels convolutional architecture

A lot of effort is currently made to provide methods to analyze and understand deep neural network impressive performances for tasks such as image or text classification. These methods are mainly based on visualizing the important input features taken into account by the network to build a decision. However these techniques, let us cite LIME, SHAP, Grad-CAM, or TDS, require extra effort to interpret the visualization with respect to expert knowledge. In this paper, we propose a novel approach to inspect the hidden layers of a fitted CNN in order to extract interpretable linguistic objects from texts exploiting classification process. In particular, we detail a weighted extension of the Text Deconvolution Saliency (wTDS) measure which can be used to highlight the relevant features used by the CNN to perform the classification task. We empirically demonstrate the efficiency of our approach on corpora from two different languages: English and French. On all datasets, wTDS automatically encodes complex linguistic objects based on co-occurrences and possibly on grammatical and syntax analysis.

Domaines

Intelligence artificielle [cs.AI] Traitement du texte et du document

Fichier principal

wTDS_HAL.pdf (980.35 Ko)

Origine : Fichiers produits par l'(les) auteur(s)

Marco Corneli : Connectez-vous pour contacter le contributeur

https://hal.science/hal-03142170

Soumis le : lundi 15 février 2021-18:45:28

Dernière modification le : lundi 11 mars 2024-15:14:04

Archivage à long terme le : dimanche 16 mai 2021-20:00:58

Dates et versions

hal-03142170 , version 1 (15-02-2021)

Identifiants

HAL Id : hal-03142170 , version 1

Citer

Laurent Vanni, Marco Corneli, Damon Mayaffre, Frédéric Precioso. From text saliency to linguistic objects: learning linguistic interpretable markers with a multi-channels convolutional architecture. 2021. ⟨hal-03142170⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

CNRS INRIA I3S DIEUDONNE INRIA2 UNIV-COTEDAZUR

78 Consultations

60 Téléchargements