Challenging Issues and Similarity Measures for Web Document Clustering

Iosr Journals, S. Mahalakshmi · Figshare · 2015

Web itself contains a large amount of documents available in electronic form. The available documents are in various forms and the information in them is not in organized form. The lack of organization of materials in the WWW motivates people to automatically manage the huge amount of information. Text- mining refers generally to the process of extracting interesting and non-trivial information and knowledge from unstructured text. Text mining framework contains Information Retrieval, Information Extraction, Information Mining and Interpretation. During Information Retrieval, so many web documents are retrieved. In that how we can find out similar documents among retrieved? This paper deals with the challenging issues and similarity measures for web document clustering .

Read the paper · More papers on PaperTik