An Approach for Document Pre-processing and K Means Algorithm Implementation

S Gowtham, Mausumi Goswami, Karthika K Balachandran, Bipul Syam Purkayastha · 2014

The web mining is a cutting edge technology, which includes information gathering and classification of information over web. This paper puts forth the concepts of document pre-processing, which is achieved by extraction of keywords from the documents fetched from the web, processing it and generating a term-document matrix, TF-IDF and the different approaches of TF-IDF (term frequency Inverse document frequency) for each respective document. The last step is the clustering of these results through K Means algorithm, by comparing the performance of each approach used. The algorithm is realized on an X64 architecture and coded on Java and Matlab platform. The results are tabulated.

Read the paper · More papers on PaperTik