Feature Selection For Text Categorisation Using Self-organising Map
Pensiri Manomaisupat, K. Abmad · 2006
The categorisation of documents in large diverse collections poses a keen problem. The choice of a vector that may represent a document collection, and categories of documents within, is still an art form. We describe a study where four different types of term occurrence and document frequency metrices have been used with varying levels of success measured by classification accuracy statistics and average quantization error; TFIDF and its variant, term relevance, have been used together with a metric based on contrastive linguistics and another uses a finely-classified terminology data base. A novel method of term representation has been used - each element of the vector corresponds to the absence/presence of a set terms colocated within the element on the basis of frequency. In addition, we have defined a new baseline for comparison - a randomly selected set of terms for constructing a representative vector from within the collection. Categorisation was performed using the classic self-organising maps. We confirm that there is an optimum size of the input vector-c.100-200 terms- exists for each of the term-occurrence/document frequency metrices, and there appears to be a saturation point beyond that optimal limit