Frequent Itemset-based Text Clustering Approach to Cluster Ranked Documents
Snehalata Nandanwar, Geetanjali Kale, Sheetal S. Sonawane · IOSR Journal of Computer Engineering · 2014
In most of the search engines documents are retrieved and ranked on the basis of relevance.They are not necessarily ranked on the basis of similarity between the query and the respective document.The ranked documents are in the form of a list.So there is a need to rank the retrieved documents.We implement Language model that is efficient to retrieve the text documents satisfying given query.The ranking of documents is based on the similarity coefficient calculated using language model.To cluster the ranked retrieved documents, we use FITC (Frequent Item set-based text Clustering) algorithm.The algorithm partitions the documents and returns the cluster in each partition.The algorithm identifies clusters with no overlap.The system accepts user request and returns all the relevant documents partitioned in the form of clusters which satisfy the query.The results of applying the algorithm to document retrieval demonstrate that the algorithm identifies non-overlapping clusters and is therefore of widespread use in many of the search engines.