A semantic-based text classification system
Abdullah Bawakid, Mourad Oussalah · 2010
This paper presents a system that performs automatic semantic-based text categorization. Using Princeton WordNet, a series of induced methods were implemented that extract semantic features from text and utilize them to decide how similar a document is to different topics. In addition, a bag-of-words method incorporating no knowledge from WordNet is implemented in the system as a basis to compare different WordNet-based approaches. This paper describes the system and reports on a simple analysis performed to evaluate the different implemented methods. At the end, a discussion on the limitations of this study and the future work to optimize the system is presented.