Web Spam Hunting @ Budapest

Dávid Siklósi, András A. Benczúr, Zsolt Fekete, Miklós Kurucz, István Bíró, Attila Pereszlényi, Simon Rácz, Adrienn Szabó, Jácint Szabó · 2008

We use a combination, in the expected order of their strength, of the following classificators: SVM over tf.idf, an augmented set of the public statistical spam features, graph stacking and text classification by latent Dirichlet allocation and compression, the latter two only used in our second submission. 1. THE METHOD We split features into related sets and for each we use the best fitting classifier. These classifiers are then combined by random forest, a method that, in our crossvalidation experiment, outperformed logistic regression suggested by [6]. We used the classifier implementations of the machine learning toolkit Weka [8]. Our expected results obtained by crossvalidation over the training data are shown in [3]. Graph stacking, a methodology used with success for Web

Read the paper · More papers on PaperTik