On Clustering and Evaluation of Narrow Domain Short-Test Corpora
DAVID EDUARDO PINTO AVENDAÑO · 2008
In this Ph.D. thesis we investigate the problem of clustering a particular set of documents namely narrow domain short texts.To achieve this goal, we have analysed datasets and clustering methods.Moreover, we have introduced some corpus evaluation measures, term selection techniques and clustering validity measures in order to study the following problems:1. To determine the relative hardness of a corpus to be clustered and to study some of its features such as shortness, domain broadness, stylometry, class imbalance and structure.2. To improve the state of the art of clustering narrow domain short-text corpora.The research work we have carried out is partially focused on "short-text clustering".We consider this issue to be quite relevant, given the current and future way people use "small-language" (e.g.blogs, snippets, news and text-message generation such as email or chat).Moreover, we study the domain broadness of corpora.A corpus may be considered to be narrow or wide domain if the level of the document vocabulary overlapping is high or low, respectively.In the categorization task, it is very difficult to deal with narrow domain corpora such as scientific papers, technical reports, patents, etc.The aim of this research work is to study possible strategies to tackle the following two problems: a) the low frequencies of vocabulary terms in short texts, and b) the high vocabulary overlapping associated to narrow domains.Each problem alone is challenging enough, however, dealing with narrow domain short texts increases the complexity of the problem significantly.The clustering of scientific abstracts is even more difficult than the clustering of narrow domain short-text corpora.The reason is that texts belonging to scientific vii viii