Automatically Identifying Pseudepigraphic Texts
Moshe Koppel, Shachar Seidman · 2013
The identification of pseudepigraphic textstexts not written by the authors to which they are attributedhas important historical, forensic and commercial applications.We introduce an unsupervised technique for identifying pseudepigrapha.The idea is to identify textual outliers in a corpus based on the pairwise similarities of all documents in the corpus.The crucial point is that document similarity not be measured in any of the standard ways but rather be based on the output of a recently introduced algorithm for authorship verification.The proposed method strongly outperforms existing techniques in systematic experiments on a blog corpus.