Using Visual Content-based Analysis with Textual and Structural Analysis for Improving Web Filtering

Mohamed Hammami 1 Liming Chen 1 Youssef Chahir 2
2 Equipe Image - Laboratoire GREYC - UMR6072
GREYC - Groupe de Recherche en Informatique, Image, Automatique et Instrumentation de Caen
Abstract : Along with the ever growingWeb is the proliferation of objectionable content, such as sex, violence, racism, etc. We need efficient tools for classifying and filtering undesirable web content. In this paper, we investigate this problem through WebGuard, our automatic machine learning based pornographic website classification and filtering system. Facing the Internet more and more visual and multimedia as exemplified by pornographic websites, we focus here our attention on the use of skin color related visual content based analysis along with textual and structural content based analysis for improving pornographic website filtering. While the most commercial filtering products on the marketplace are mainly based on textual content-based analysis such as indicative keywords detection or manually collected black list checking, the originality of our work resides on the addition of structural and visual content-based analysis to the classical textual content-based analysis along with several major-data mining techniques for learning and classifying. Experimented on a testbed of 400 websites including 200 adult sites and 200 non pornographic ones, WebGuard, our Web filtering engine scored a 96.1% classification accuracy rate when only textual and structural content based analysis are used, and 97.4% classification accuracy rate when skin color related visual content based analysis is driven in addition. Further experiments on a black list of 12 311 adult websites manually collected and classified by the French Ministry of Education showed that WebGuard scored 87.82% classification accuracy rate when using only textual and structural content-based analysis, and 95.62% classification accuracy rate when the visual content-based analysis is driven in addition. The basic framework of WebGuard can apply to other categorization problems of websites which combine, as most of them do today, textual and visual content.
Complete list of metadatas

https://hal.archives-ouvertes.fr/hal-00822231
Contributor : Yvain Queau <>
Submitted on : Tuesday, May 14, 2013 - 12:29:59 PM
Last modification on : Tuesday, November 19, 2019 - 2:38:41 AM

Links full text

Identifiers

Citation

Mohamed Hammami, Liming Chen, Youssef Chahir. Using Visual Content-based Analysis with Textual and Structural Analysis for Improving Web Filtering. International Journal of web information systems, Emerald, 2005, 1 (4), pp.241-254. ⟨10.1108/17440080580000096⟩. ⟨hal-00822231⟩

Share

Metrics

Record views

298