Probabilistic and Possibilistic Language Models Based on the World Wide Web

Stanislas Oger; Vladimir Popescu; Georges Linarès

Communication Dans Un Congrès Année : 2009

Probabilistic and Possibilistic Language Models Based on the World Wide Web

(1) , (1) , (1)

Stanislas Oger

Fonction : Auteur
PersonId : 770872
IdRef : 176527176

Laboratoire Informatique d'Avignon

Vladimir Popescu

Fonction : Auteur

Laboratoire Informatique d'Avignon

Georges Linarès

Fonction : Auteur
PersonId : 4977
IdHAL : georges-linares
IdRef : 079368794

Laboratoire Informatique d'Avignon

Résumé

Usually, language models are built either from a closed corpus, or by using World Wide Web retrieved documents, which are considered as a closed corpus themselves. In this paper we propose several other ways, more adapted to the nature of the Web, of using this resource for language modeling. We first start by improving an approach consisting in estimating n-gram probabilities from Web search engine statistics. Then, we propose a new way of considering the information extracted from the Web in a probabilistic framework. Then, we also propose to rely on Possibility Theory for effectively using this kind of information. We compare these two approaches on two automatic speech recognition tasks: (i) transcribing broadcast news data, and (ii) transcribing domain-specific data, concerning surgical operation film comments. We show that the two approaches are effective in different situations. Index Terms: language modeling, World Wide Web, possibility measure, automatic speech recognition

Domaines

Informatique [cs]

bibliothèque Universitaire Déposants HAL-Avignon : Connectez-vous pour contacter le contributeur

https://hal.science/hal-01319863

Soumis le : lundi 23 mai 2016-10:27:40

Dernière modification le : mardi 22 mars 2022-14:40:01

Dates et versions

hal-01319863 , version 1 (23-05-2016)

Identifiants

HAL Id : hal-01319863 , version 1

Citer

Stanislas Oger, Vladimir Popescu, Georges Linarès. Probabilistic and Possibilistic Language Models Based on the World Wide Web. INTERSPEECH, Sep 2009, Brighton, United Kingdom. ⟨hal-01319863⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

UNIV-AVIGNON LIA

27 Consultations

0 Téléchargements

Probabilistic and Possibilistic Language Models Based on the World Wide Web

Résumé

Domaines

Dates et versions

Identifiants

Citer

Exporter

Collections

Partager