On the development of a method for categorizing domain-specific text snippets
M.M. El Maslouhi · Utrecht University Repository (Utrecht University) · 2013
This thesis focuses on the automatic categorization of domain-specific text snippets. The main research question was defined as “How can a categorization method be developed that is able to categorize domain–specific text snippets?”. The main contribution of this thesis is the developed method that addresses this issue. The method was designed by identifying the issues that accompany the main research question. The main issues with text snippets is that they lack contextual perspectives and suffer from data-sparseness. The literature study elaborates the state-of-the-art from the solution domains. The concepts from these solution domains provide the scientific grounding of the method. The main concepts were retrieved from the information retrieval domain and information extraction domain. The proposed method introduces a taxonomy to enrich the text snippets with contextual data, feature selection is used to identify informative terms and finally a probabilistic approach is applied to determine the most likely category based on a training set. To validate the methods’ effectiveness measures were employed to validate the performance of the classifier. The main effectiveness measure was the Fscore which illustrates the overall effectiveness of the classifier. The validation showed encouraging results in the sense that there are conclusive causal relationships between the components of the method and the results. The results show that the taxonomy and the feature selection techniques have a profound effect on the performance of the classifier.