Preprocessing Techniques in Text Categorization
Pritam C. Gaigole, Leena H. Patil, P. M Chaudhari · 2013
Bulk data is generated in the era ofInformation Technology. If it is not stored in aproperly systematic manner then the generated datacannot be reused. This is because navigation becomes if not impossible, certainly very difficult. The data generated is to analyze so as to maximizethe benefits, for intelligent decision making. Textcategorization is an important and extensively studiedproblem in machine learning. The basic phases in textcategorization include preprocessing features, extractingrelevant features against the features in a database, andfinally categorizing a set of documents into predefinedcategories. Most of the researches in text categorization arefocusing more on the development of algorithms andcomputer techniques.