Towards Lexical Encoding of Multi-Word Expressions in Spanish Dialects - Archive ouverte HAL Accéder directement au contenu
Communication Dans Un Congrès Année : 2016

Towards Lexical Encoding of Multi-Word Expressions in Spanish Dialects

Résumé

This paper describes a pilot study in lexical encoding of multi-word expressions (MWEs) in 4 Latin American dialects of Spanish: Costa Rican, Colombian, Mexican and Peruvian. We describe the variability of MWE usage across dialects. We adapt an existing data model to a dialect-aware encoding, so as to represent dialect-related specificities, while avoiding redundancy of the data common for all dialects. A dozen of linguistic properties of MWEs can be expressed in this model, both on the level of a whole MWE and of its individual components. We describe the resulting lexical resource containing several dozens of MWEs in four dialects and we propose a method for constructing a web corpus as a support for crowdsourcing examples of MWE occurrences. The resource is available under an open license and paves the way towards a large-scale dialect-aware language resource construction, which should prove useful in both traditional and novel NLP applications.
Fichier principal
Vignette du fichier
65_Paper.pdf (1.47 Mo) Télécharger le fichier
Origine : Fichiers éditeurs autorisés sur une archive ouverte
Loading...

Dates et versions

hal-01505049 , version 1 (10-04-2017)

Identifiants

  • HAL Id : hal-01505049 , version 1

Citer

Diana Bogantes, Eric Rodríguez, Alejandro Arauco, Alejandro Rodríguez, Agata Savary. Towards Lexical Encoding of Multi-Word Expressions in Spanish Dialects. Tenth International Conference on Language Resources and Evaluation (LREC 2016), May 2016, Portorož, Slovenia. ⟨hal-01505049⟩
103 Consultations
50 Téléchargements

Partager

Gmail Facebook X LinkedIn More