Kernel-based text categorisation

Radwan Jalam, Olivier Teytaud · 2002

Presents some techniques in text categorization. New algorithms, in particular a new support vector machine kernel for text categorization, are developed and compared to usual techniques. This kernel leads to a more natural space for elaborating separations than the euclidian space of frequencies or even in verse frequencies, as the distance in this space is the most usual pseudo-distance between distributions. We give an application to the recognition of the author of a text, and put into relief that our kernel could be used for any classification of distributions. We experimentally discuss the efficiency of our algorithms, depending on the precision of the estimation of frequencies, and the possibility of building statistical bounds on the error. All our experiments are made on underconstrained problems.

Read the paper · More papers on PaperTik