A Study for Evaluating the Importance of Various Parts of Speech (POS) for Information Retrieval (IR)

Chirag Suresh Shah · 2002

Traditionally in the vector space model of document representation for various IR (Information Retrieval) tasks, all the content words are used without considering their individual significance in the language. Such methods treat a document as a bag-of-words and do not exploit any language related information. It is obvious that considering such information in representing the documents can help in improving the performance of various IR tasks, but how to obtain this information is considered to be difficult. One of the information that can be important is the knowledge about the role of various parts-of-speech (POS). Although importance of various POS is very subjective and depends on the application as well as the domain under consideration, it can be very useful to evaluate their importance even in a general setup. In this paper we present a study to understand this importance. We first generate the document vectors using particular POS. We then evaluate how good is this representation. This is done by measuring the information content provided by document vectors. This information is then used to reconstruct the document vectors. In order to show that these document vectors are better than those of generated by traditional methods, we consider text classification application. We show some improvement in classification accuracy, but more importantly, we demonstrate the consistency in the results and a step toward a new and promising direction for using semantics for IR tasks.

Read the paper · More papers on PaperTik