An Efficient Inclusive Similarity Based Clustering (ISC) Algorithm for Big Data

J. Sangeetha, V. Sinthu Janita Prakash · 2017

In recent days, the big data opinion mining is considered as an important research area because, the size of the user reviews increase in petabytes. Opinion classification is the process of identifying the positive, negative and neutral opinion from the text. The traditional data mining software tools find it difficult to manage the reviews because of its size. Hence, to provide an efficient big data opinion mining, an efficient Inclusive Similarity based Clustering (ISC) algorithm is proposed. During pre-processing the data is cleaned with Parts of Speech (PoS) tagger and sliced by using the proposed Threshold based Data Partitioning (TDP) algorithm. After pre-processing the dataset, the inclusive similarity algorithm computes the similarity between the consecutive reviews and constructs an adjacency matrix. The proposed ISC algorithm exploits the adjacency matrix and merges the clusters into a single cluster. The performance of the proposed method is validated with the existing algorithms for the metrics such as time consumption, memory utility and accuracy. The comparison results prove that the proposed ISC algorithm provides optimal results for all the metrics than the existing algorithms.

Read the paper · More papers on PaperTik