A REVIEW PAPER ON ALGORITHMS USED FOR TEXT CLASSIFICATION

Sukhjit Singh Sehra · 2013

The textual revolution has seen a tremendous change in the availability of online information. Finding information for just about any need has never been more automatic. Text classification (also known as text categorization or topic spotting) is the task of automatically sorting a set of documents into categories from a predefined set. This task has several applications, including automated indexing of scientific articles, filing patents into patent directories, selective dissemination of information to information consumers, automated population of hierarchical catalogues of Web resources, spam filtering, identification of document genre. Automated text classification is attractive because it frees organizations from the need of manually organizing document bases, which can be too expensive, or simply not feasible given the time constraints of the application or the number of documents involved. The accuracy of modern text classification systems rivals that of trained human professionals, thanks to a combination of information retrieval (IR) technology and machine learning (ML) technology. The aim of this paper is to highlight the important algorithms that are employed in text documents classification, while at the same time making awareness of some of the interesting challenges that remain to be solved.

Read the paper · More papers on PaperTik