Uncovering Languages from written documents
Nikitas Ν. Karanikolas, Panagiotis Ouranos · 2014
Language identification is the process of understanding what is the language used in a document. There are non-computational (manual) approaches and also computational ones. More problematic are cases where the input text is composed of several languages. This is a common situation on Web documents. Our approach is a computational language identification process which can address multiple-language documents from the Web. It elaborates the used single--byte alphabet (in cases of codepage based texts), the used subset of Unicode characters (in cases of multiple--byte encoded texts), the prevalence of certain function words, n--grams and other more granular symbol manipulations. We discuss the methods and evaluate the system with resources (documents) with a mixture of languages (Bulgarian, Byelorussian, Russian, Greek, Serbian, etc).