Mistake-driven learning with thesaurus for text categorization
Takefumi Yamazaki, Ido Dagan · 1997
This paper extends the mistake-driven learner WINNOW to better utilize thesauri for text categorization. In our method not only words but also semantic categories given by the thesaurus are used as features in a classifier. New filtering and disambiguation methods are used as pre-processing to solve the problems caused by the use of the thesaurus. In order to verify our methods, we test a large body of tagged Japanese newspaper articles created by RWCP 1 . Experimental results show that WINNOW with thesauri attains high accuracy and that the proposed filtering and disambiguation methods also contribute to the improved accuracy. 1 Introduction It has become essential to automatically organize or classify the huge amount of electric documents if we are to handle the "information flood". Obviously, information retrieval tasks such as text categorization and routing are becoming more and more important. These tasks can be considered as classification problems. In text categorization, g...