Dynamic Similarity Threshold in Authorship Verification: Evidence from Classical Arabic
Hossam Ahmed · Procedia Computer Science · 2017
Many Authorship Verification Machine Learning-based algorithms rely on establishing a similarity threshold θ between a candidate text and known texts in terms of one or more linguistic features. Documents that score below that threshold are rejected as not written by the same author. Current definitions of θ rely on both negative and positive training input. An algorithm that relies exclusively on positive training data, and dynamically calculates θ for each verification problem performs with good accuracy, tested using a training and evaluation corpus from Classical Arabic.