Skip to Main content Skip to Navigation
Journal articles

The NLP4NLP Corpus (I): 50 Years of Publication, Collaboration and Citation in Speech and Language Processing

Abstract : This paper introduces the NLP4NLP corpus, which contains articles published in 34 major conferences and journals in the field of speech and natural language processing over a period of 50 years (1965–2015), comprising 65,000 documents, gathering 50,000 authors, including 325,000 references and representing ~270 million words. Most of these publications are in English, some are in French, German, or Russian. Some are open access, others have been provided by the publishers. In order to constitute and analyze this corpus several tools have been used or developed. Many of them use Natural Language Processing methods that have been published in the corpus, hence its name. The paper presents the corpus and some findings regarding its content (evolution over time of the number of articles and authors, collaborations between authors, citations between papers and authors), in the context of a global or comparative analysis between sources. Numerous manual corrections were necessary, which demonstrated the importance of establishing standards for uniquely identifying authors, articles, or publications.
Complete list of metadatas

https://hal.archives-ouvertes.fr/hal-02413751
Contributor : Limsi Publications <>
Submitted on : Monday, December 16, 2019 - 12:42:09 PM
Last modification on : Monday, February 10, 2020 - 6:14:09 PM

Identifiers

  • HAL Id : hal-02413751, version 1

Citation

Joseph Mariani, Gil Francopoulo, Patrick Paroubek. The NLP4NLP Corpus (I): 50 Years of Publication, Collaboration and Citation in Speech and Language Processing. Frontiers in Research Metrics and Analytics, Frontiers Media, 2019, 3, pp.1-30. ⟨hal-02413751⟩

Share

Metrics

Record views

11