Algorithm for Punjabi Text Classification
Nidhi Nidhi, Vishal Gupta · International Journal of Computer Applications · 2012
Text Mining is a field that extracts hidden, not yet discovered, useful information from the text document according to user’s query. And Text Classification is one of the text mining tasks to manage the information efficiently, by classifying the documents into classes using classification algorithms. Any text classification method uses a set of features to characterize each text document, where these features should be relevant to the task at hand. Not much work has been done for Punjabi text classification. Adequate annotated corpora are not yet available in Punjabi. This paper introduces preprocessing techniques, features selection methods for Punjabi and classification algorithm to classify the Punjabi Text documents. General Terms Natural language processing, Text classification