Classifying web pages by content

Dan J. Smith · 1999

This paper describes a classification strategy for multimedia documents and reviews the prospects for detecting and filtering documents, such as Web pages, that may be pornographic. We examine several colour filtering algorithms with a view to producing a reliable skin filter. The results, very simple features extracted from an image-only database containing around two-thousand hand-labelled images, are surprisingly good. When the image results are combined with a simple text analysis scheme we are able to achieve a very accurate classification.

Read the paper · More papers on PaperTik