The NLP4NLP Corpus (I): 50 Years of Publication, Collaboration and Citation in Speech and Language Processing - Archive ouverte HAL Accéder directement au contenu
Article Dans Une Revue Frontiers in Research Metrics and Analytics Année : 2019

The NLP4NLP Corpus (I): 50 Years of Publication, Collaboration and Citation in Speech and Language Processing

Résumé

This paper introduces the NLP4NLP corpus, which contains articles published in 34 major conferences and journals in the field of speech and natural language processing over a period of 50 years (1965–2015), comprising 65,000 documents, gathering 50,000 authors, including 325,000 references and representing ~270 million words. Most of these publications are in English, some are in French, German, or Russian. Some are open access, others have been provided by the publishers. In order to constitute and analyze this corpus several tools have been used or developed. Many of them use Natural Language Processing methods that have been published in the corpus, hence its name. The paper presents the corpus and some findings regarding its content (evolution over time of the number of articles and authors, collaborations between authors, citations between papers and authors), in the context of a global or comparative analysis between sources. Numerous manual corrections were necessary, which demonstrated the importance of establishing standards for uniquely identifying authors, articles, or publications.

Dates et versions

hal-02413751 , version 1 (16-12-2019)

Licence

Paternité

Identifiants

Citer

Joseph J Mariani, Gil Francopoulo, Patrick Paroubek. The NLP4NLP Corpus (I): 50 Years of Publication, Collaboration and Citation in Speech and Language Processing. Frontiers in Research Metrics and Analytics, 2019, 3, pp.1-30. ⟨10.3389/frma.2018.00036⟩. ⟨hal-02413751⟩
72 Consultations
0 Téléchargements

Altmetric

Partager

Gmail Facebook X LinkedIn More