Significant Feature Clustering
John S. Whissell · 2006
I hereby declare that I am the sole author of this thesis. This is a true copy of the thesis, including any required final revisions, as accepted by my examiners. I understand that my thesis may be made electronically available to the public. iii In this thesis, we present a new clustering algorithm we call Significance Feature Clustering, which is designed to cluster text documents. Its central premise is the mapping of raw frequency count vectors to discrete-valued significance vectors which contain values of-1, 0, or 1. These values represent whether a word is significantly negative, neutral, or significantly positive, respectively. Initially, standard tf-idf vectors are computed from raw frequency vectors, then these tf-idf vectors are transformed to significance vectors using a parameter α, where α controls the mapping-1, 0, or 1 for each vector entry. SFC clusters agglomeratively, with each document’s significance vector representing a cluster of size one containing just the document, and iteratively merges the two clusters that exhibit the most similar