Automated analysis of Bangla poetry for classification and poet identification

Geetanjali Rakshit, Anupam Ghosh, Pushpak Bhattacharyya, Gholamreza Haffari · 2015

Computational analysis of poetry is a challenging and interesting task in NLP. Human expertise on stylistics and aesthetics of poetry is generally expensive and scarce. In this work, we delve into the data to automati-cally extract stylistic and linguistic in-formation which are useful for anal-ysis and comparison of poems. We make use of semantic (word) features to perform subject-based classification of Bangla poems, and various stylistic as well as semantic features for poet iden-tification. We have used a Multiclass SVM classifier to classify Tagore’s col-lection of poetry into four categories: devotional, love, nature and national-ism. We identified the most useful word features for each category of po-ems. The overall accuracy of the classi-fier was 56.8%, and the analysis led us to conclude that for poetry classifica-tion, word features alone do not suffice, due to allusions often being used as a poetic device. We, next, used these fea-tures along with stylistic features (syn-tactic, orthographic and phonemic), for poet identification on a dataset of po-ems from four poets and achieved a performance of 92.3 % using a Multi-class SVM classifier. While content-based and stylometric analysis of prose in Bangla has been done in the past, this is a first such attempt for poetry. 1

Read the paper · More papers on PaperTik