Detection of Foreign Entities in Native Text Using N-gram Based Cumulative Frequency Addition

Bashir Ahmed, Sung-Hyuk Cha, Charles C. Tappert · 2005

This paper describes a logarithmic version of the conventional Naive Bayesian N-gram-based, textclassification algorithm that we name Cumulative Frequency Addition (CFA) and its application in three tasks: language identification, nationality identification from names, and detection of foreign words in base text. The new CFA technique is 3-10 times faster than N-gram based rank-order statistical classifiers. In the language identification task CFA yields 100% accuracy on string sizes greater than 150 characters. In the name-tonationality task, it yields 86% accuracy on a 14 country database and 96% on a 7 country database within the top three choices. Finally, in the task of detecting foreign words it yields 66.9% accuracy. This is the first study to apply natural language processing techniques to such tasks as name identification and foreign word detection.

Read the paper · More papers on PaperTik