Performance evaluation of categorizing technical support requests using advanced K-Means algorithm

Mubina A. Nadaf, Sachin Patil · 2015

Technical support service providers receive thousands of customer queries daily. Traditionally, such organizations discard the data due to lack of storage capacity. However, value of storing such data is needed for the better results of analysis and to improve the closure rate of the daily customer queries. Data mining is the process of finding important and meaningful information, patterns through the large amount of data. Clustering is used as one of the best concept for data analysis, using machine learning approach with mathematical and statistical methods. Cluster analysis is widely applicable for practical applications in emerging trends in data mining. Analysis of clustering algorithms such as K-Means, Dirichlet, Fuzzy K-Means Canopy algorithms is done by means of the practical approach, in this research work. Performance of algorithm is observed based on the execution or computational time and results are compared with each of these algorithms. This paper proposes the streaming K-Means algorithm which resolves the queries as it arrives and analyses the data. Cosine distance measure plays an important role in clustering dataset. Sum of Square error is measured to check the quality of the cluster.

Read the paper · More papers on PaperTik