Distinguishing between True and False Stories using various Linguistic Features

Yaakov HaCohen‐Kerner, Rakefet Dilmon, Shimon Friedlich, Daniel Nisim Cohen · 2015

This paper analyzes what linguistic features differentiate true and false stories written in Hebrew. To do so, we have defined four feature sets containing 145 features: POS-tags, quantitative, repetition, and special expressions. The examined corpus contains stories that were composed by 48 native Hebrew speakers who were asked to tell both false and true stories. Classification experiments on all possible combinations of these four feature sets using five supervised machine learning methods have been applied. The Part of Speech (POS) set was superior to all others and has been found as a key component. The best accuracy result (89.6%) has been achieved by a combination of sixteen POS-tags and one quantitative feature. 1

Read the paper · More papers on PaperTik