Punjabi Poetry Classification
Jasleen Kaur, Jatinderkumar R. Saini · 2017
Literature of country represents the prosperity of that country. India, being a multilingual country, is having a rich heritage and literature. In order to retrieve literature pieces easily, it must be classified. In this research work, vocabulary-content based classification of Punjabi poetry is done. 4 different poetry categories are populated with 240 poetries (with 60 poems in each category). These 240 poetry documents are passed through typical NLP text classification phases like Sentence Splitting, Tokenization and Bag-of-Words (BOW) representation, finally yielding to their Vector Space Model (VSM) representation. Total 9867 unique words extracted from last step are used for building the different machine learning models. For the first time in research community, 10 different machine learning algorithms are trained and tested for any Indian language, using weka, with an aim to find the most suitable algorithm. Results for Punjabi poetry classification revealed that 4 machine learning algorithms namely, Hyperpipes (HP), K- nearest neighbour (KNN), Naive Bayes (NB) and Support Vector Machine (SVM) with an accuracy of 50.63 %, 52.92 %, 52.75 % and 58.79 % respectively, outperformed all other machine learning algorithms under the test.