Scalable Clustering Using Rank Based Pre-processing Technique for Mixed Data Sets Using Enhanced Rock Algorithm
P. Parameswari, J. Abdul Samath, S. Mohana Saranya, Sri Ramakrishna · 2015
The current requirements to cluster real world data sets are scalability, ability to handle any kind of data like categorical and numerical . It should also have the capability to handle noisy and missing data. Traditional algorithm can cluster categorical or numerical data but not the both. In general it is tedious to cluster mixed data types but it gives us best clusters with more accurate results. Another important factor that affects the quality of clusters are preprocessing techniques. In order to meet out the current requirement we proposed a clustering methodology that helps to enhance the performance of ROCK clustering algorithm which is scalable. This approach has two process (1) Numerical attributes are converted in to categorical, missing values are filled by using a rank based method (2) Clustering takes place using ROCK algorithm. These approaches are combined together and known as EROCK algorithm. Experimental results obtained by this methodology are compared with EM and CLOPE algorithms. It shows that our new methodology performs well for real world data sets and found it is very effective.