A concept driven document clustering using WordNet
Sujata R. Kolhe, Dr.Sudhir Deoraoji Sawarkar · 2017
Recently Information Technology is used extensively for wide range of application for example solutions enabled through e-commerce to different web based information system. This usage has lead to development of large textual data base. Mostly this information data is stored in unstructured text. This large data developed has lead to the need of its systematic clustering for easy data retrieval organization and summarization, typically called as data mining. Many techniques such as document clustering, web based result clustering and information visualization are developed by the researchers based on various algorithms such as k-means, k-medoid, Bisecting k-means, Suffix Tree Clustering (STC), Hierarchical clustering, Lingo and nature inspired optimization algorithm like clustering using Ant colony optimization, Fireflies and so on. Although these techniques are widely used, these are suffered from inconsistencies in cluster content and inconsistencies in cluster description. Therefore this paper presents a novel approach termed as `Semantic Lingo' based on concept of the text data. This concept driven approach is executed on Wordnet. The developed approach identifies the predominant notion and automatically generate clusters based on these notion. The standard available dataset ODP, AMBIENT are used in the paper for validation of the proposed concept. The data clusters obtained from the proposed concept has shown high precision level and recall.