Review of Web Document Clustering algorithms

Sanjib Kumar Sahu, Shalini Srivastava · 2016

Users of search engines are fond of accurate and fast results. Web is now overloaded with lots of pages or documents dealing with same topic. Thus they are often forced to surf through the large and irrelevant set of results. This has forced the IR community to explore such document clustering techniques capable of providing fast and accurate results. Even after being a very effective solution, clustering is yet not deployed on the major search engines. This paper will articulate the requirements of Web Document Clustering and reports on the clustering methods belonging in this domain. The focus of ours is on; these methods create their clusters based on the characters or individual terms rather than showcasing them as a single phrase with a meaning and sequence of words. Paper will be reporting general term based and phrase based techniques and will provide conclusion based on their individual efficiency to work with their key methods.

Read the paper · More papers on PaperTik