Determining of discriminative blog size for authorship attribution on the Turkish texts

Pelin Canbay, Hayri Sever, Ebru Akçapınar Sezer · 2018

Although many features and methods are used to extract information from a text about its author, a standard method and feature set could not be presented in this area. The fact that different types of texts are produced continuously in the electronic environment has made this process even more difficult. Authorship attribution, which is interested to find the author of the anonymous text, is a common branch of forensic science, computer science, and linguistics. This study focuses on answering the question of what is the discriminative and satisfying text size for authorship attribution studies. The study conducted on the Turkish blog writings and aimed at providing a standard solution step in this area. As a result of the many experiments, short texts from 500 words are seemed inappropriate to find meaningful results in authorship attribution studies.

Read the paper · More papers on PaperTik