Using Semantics in Document Representation: A Lexical Chain Approach

Dinakar Jayarajan · 2009

Automatic classification and clustering are two of the most common operations performed on text documents. Numerous algorithms have been proposed for this and invariably, all of these algorithms use some variation of the vector space model to represent the documents. Traditionally, the Bag of Words (BoW) representation is used to model the documents in a vector space. The BoW scheme is a simple and popular scheme, but it suffers from numerous drawbacks. The chief among them is that the feature vectors generated using BoW results in very large dimensional vectors. This creates problems with most machine learning algorithms where the high dimensionality severely affects the performance. This fact is manifested in the current thinking in the machine learning community, that some sort of a dimensionality reduction is a beneficial preprocessing step for applying machine learning algorithms on textual data. BoW also fails to capture the semantics contained in the documents properly. In this thesis, we present a new representation for documents based on lexical chains. This representation addresses both the problems with BoW it achieves a significant reduction in the dimensionality and captures some of the semantics present in the data. We present an improved algorithm to compute lexical chains and generate feature vectors using these chains. We evaluate our approach using datasets derived from the 20 Newsgroups corpus. This corpus is a collection of 18,941 documents across 20 Usenet groups related to topics such as computers, sports, vehicles, religion, etc. We compare the performance of the lexical chain based document features against the BoW features on two machine learning tasks clustering and classification. We also present a preliminary algorithm for soft clustering using the lexical chains based representation, to facilitate topic detection from a document corpus.

Read the paper · More papers on PaperTik