Addressing challenges in automatic Language Identification of Romanized Text

K. Pavan, Niket Tandon, Vasudeva Varma · 2013

Due to the diversity of documents on web, language identification is a vital task for web search engines during crawling and indexing of web documents. Among the current challenges in language-identification, the unsettled problem remains identifying Romanized text language. The challenge in Romanized text is the variations in word spellings and sounds in different dialects. We propose a Romanized text language identification system (RoLI) that addresses these challenges. RoLI uses an n-gram based approach and also exploits sound based similarity of words. RoLI does not rely on language intensive resources and is robust to Multilingual text. We focus on five Indian languages: Hindi, Telugu, Tamil, Kannada and Malayalam. Over the five languages, we achieve an average accuracy of 98.3%, despite the spelling variations as well as sound variations in Indian languages. 1

Read the paper · More papers on PaperTik