Text classification model framework based on social annotation quality
Xiwu Gu · Journal of Computer Applications · 2012
Social annotation is a form of folksonomy,which allows Web users to categorize Web resource with text tags freely.It usually implicates fundamental and valuable semantic information of Web resources.Consequently,social annotation is helpful to improve the quality of information retrieval when applied to information retrieval system.This paper investigated and proposed an improved text classification algorithm based on social annotation.Because social annotation is a kind of folksonomy and social tags are usually generated arbitrarily without any control or expertise knowledge,there has been significant variance in the quality of social tags.Under this consideration,the paper firstly proposed a quantitative approach to measure the quality of social tags by utilizing the semantic similarity between Web pages and social tags.After that,the social tags with relatively low quality were filtered out based on the quality measurement and the remained social tags with high quality were applied to extend traditional vector space model.In the extended vector space model,a Web page was represented by a vector in which the components were the words in the Web page and tags tagged to the Web page.At last,the support vector machine algorithm was employed to perform the classification task.The experimental results show that the classification result can be improved after filtering out the social tags with low quality and embedding those high quality social tags into the traditional vector space model.Compared with other classification approaches,the classification result of F1 measurement has increased by 6.2% on average when using the proposed algorithm.