STATISTICAL LANGUAGE IDENTIFICATION OF SHORT TEXTS

Fela Winkelmolen, Viviana Mascardi · 2011

Although correctly identifying the language of short texts should prove useful in a large number of applications, few satisfactory attemps are reported in the literature.In this paper we describe a Naive Bayes Classifier that performs well on very short texts, as well as the corpus that we created from movie subtitles for training it.Both the corpus and the algorithm are available under the GNU Lesser General Public License.

Read the paper · More papers on PaperTik