Language identification using Fuzzy-SVM technique

Girish Mishra, Sohan Lal Nitharwal, Sarvjeet Kaur · 2010

Language Identification is an important issue in today's multilingual world. In this paper we have analyzed Fuzzy-SVM technique for identification of romanized plaintexts of five Indian regional languages namely Hindi, Bangla, Manipuri, Urdu and Kashmiri. Distinguishing features/characteristics have been extracted from romanized plaintexts of each of these five languages and represented suitably through Fuzzy Sets on a normalized scale. These normalized feature vectors are given as input to the Support Vector Machine (SVM) based classifier. For constructing the hyperplane in a higher dimension space Guassian Radial Basis Kernal function has been used. The proposed Pattern Recognition (PR) system is independent of the dictionaries of these languages and can even identify plaintext with unknown word boundaries. This PR system (Language Identifier) can be used for automatic segregation of plain texts of these languages while analyzing intercepted, multiplexed and interleaved Speech/Data/Fax communication. The proposed method significantly improves the classification accuracy compared to the other methods even for smaller text length messages.

Read the paper · More papers on PaperTik