Tracking Web Spam with Hidden Style Similarity

Tanguy Urvoy, Thomas Lavergne, Pascal Filoche · Adversarial Information Retrieval on the Web · 2006

Automatically generated content is ubiquitous in the web: dynamic sites built using the three-tier paradigm are good examples (e.g. commercial sites, blogs and other sites pow- ered by a web authoring software), as well as less legitimous spamdexing attempts (e.g. link farms, faked directories...). Those pages built using the same generating method (tem- plate or script) share a common \look and feel that is not easily detected by common text classiflcation methods, but is more related to stylometry. In this paper, we present a (hidden) style similarity mea- sure based on extra-textual features in html source code. We also describe a method to clusterize a large collection of documents according to this measure. The clustering algo- rithm being based on flngerprints, we also give some recalls about flngerprinting. By conveniently sorting the generated clusters, one can ef- flciently track back instances of a particular automatic con- tent generation method among web pages collected using a crawler. This is particularly useful to detect pages across difierent sites sharing the same design | this is often a good hint of either spamdexing attempt or mirrored content.

Read the paper · More papers on PaperTik