Authorship Clustering using TF-IDF weighted Word-Embeddings
Lucky Agarwal, Kartik Thakral, Gaurav Bhatt, Ankush Mittal · 2019
In this paper, we propose a novel idea to solve the problem of Author Clustering which is introduced in PAN-2017 Author Identification task. First, we use word embeddings that brings semantic feature of the words along with the grammatical and syntactic information. Secondly, the documents are represented as the TF-IDF weighted sum of the embedding vectors corresponding to each word. A higher TF-IDF weight implies that the words have a stronger relationship in the documents in which they appear. Lastly, using hierarchical clustering all those documents written by each author are grouped together. Finally, we compare the proposed technique with the top submissions to PAN-2017 author clustering task. Through extensive experimentation, we show that the proposed technique outperforms various state-of-the-art models on the given task.