Comparison of Language Identification Techniques

Leonid Panich, Stefan Conrad, Martin Mauve · 2015

Many researches that analyse huge amounts of text data, eliminate only texts in English or use the datasets of texts, which language is identified. Accordingly, the language identification task for this text data is assumed to be accomplished. The purpose of the present work is to compare the language identification approaches using the datasets of tweets for the evaluation. The parameters and classifiers with the best performance for the wordand N-gram-based approaches are determined for four different datasets of tweets, that contain 19 different languages. Moreover, the approaches with the highest results are found for each of these datasets. The lists of sentences and words from the Leipzig Corpora Collection are used as the training data. The results of the present work show that for all used datasets the frequent words approach outperforms the short words approach and works with cumulative frequency addition classifier better than with other classifiers. The frequent words approach achieved the best results using 3100-3800 most frequent words. For most of the used datasets the improved graph-based N-gram approach, that utilises the natural logarithm of the counts of the N-grams, obtains the best performance. This approach shows the best results with the N-grams of the length from 3 to 5 and is used in all comparisons with the cumulative frequency addition classifier. However, for the Non-Latin dataset it is surpassed by the frequent words approach with the cumulative addition classifier and 3100 words.

Read the paper · More papers on PaperTik