Compression and stylometry for author identification

Daniel Pavelec, Luiz S. Oliveira, Edson Jose Rodrigues Justino, Francisco Dantas Nobre Neto, Leonardo Vidal Batista · 2009

In this paper we compare two different paradigms for author identification. The first one is based on compression algorithms where the entire process of defining and extracting features and training a classifier is avoided. The second paradigm, on the other hand, takes into account the classical pattern recognition framework, where linguistic features proposed by forensic experts are used to train a Support Vector Machine classifier. Comprehensive experiments performed on a database composed of 20 writers show that both strategies achieve similar performance but with an interesting degree of complementarity demonstrated through the confusion matrices. Advantages and drawback of both paradigms are also discussed.

Read the paper · More papers on PaperTik