Ontology-Based Fuzzy Semantic Clustering

Yang Cheng · 2008

Document clustering plays an important role in providing intuitive navigation and browsing mechanisms by organizing large amounts of documents into a small number of meaningful clusters. Most of the documents clustering methods were grounded in the bag of words representation to measure similarity, ignoring the semantic relationships between words that do not co-occur literally. A novel fuzzy semantic method that integrates ontology as background knowledge into the process of computing similarity between documents is proposed so as to improve the performance of documents clustering in terms of quality and efficiency. Ontology is represented as a graph-based model that reflects semantic relationship between concepts, with which a semantic similarity matrix of concepts that exploits semantic relation of the ontology is defined. Based on conceptual matrix a document can be represented to a semantic fuzzy set. Then similarity between documents is computed with fuzzy matching measure. The result of this process may make documents not similar with vector representation become similar. Maximal fuzzy spanning tree algorithm is used as a document-clustering algorithm. Finally the efficacy of our approach is demonstrated through relevant experiments.

Read the paper · More papers on PaperTik