A variant of n-gram based language-independent text categorization

Jelena Graovac · Intelligent Data Analysis · 2014

A technique for automated categorization of text documents, based on byte-level n-gram profiles and a new dissimilarity measure between profiles is presented. K nearest neighbors classifier is used. The technique is language independent. It has been

Read the paper · More papers on PaperTik