Author identification using Sequential Minimal Optimization
John Jenkins, William Nick, Kaushik Roy, Albert Esterline, Joël S. Bloch · 2016
Author identification is a substantial factor in the global economic loss due to computer-related crimes. According to the Center for Strategic and International Studies (CSIS), computer crimes or cyber-crimes cost the global economy an estimated 375 to 575 billion dollars each year [1]. Recently, various techniques have been used to improve the accuracy of author identification. In this paper, we propose combining unigram features and a variety of stylometric features that include n-grams and part-of-speech. Using a Reuters Corpus dataset of 2,500 unique articles (50 authors with 50 news articles each), we were able to effectively capture a non-topic sensitive sample. Results with the Weka machine learning software produced classification accuracies ranging from 76.08 to 84.88 percent using classification techniques such as Random Forest and Sequential Minimal Optimization (SMO). Weka also ranked and weighted the most influential feature attributes.